/teal-sea
teal-sea / zeta-labstate of record · compiled 28 Sep 2026 · revision e4945c4 · source

Library · hunts/aimo2/PREREGISTRATION.md

Preregistration: AIMO-2 legal method and controls

1,175 words · 140 lines · source

Frozen 2026-08-22, before any predictor is fitted against the robustness labels. The descriptive curation finding (reconstruct_curation.py) is not a fitted predictor and does not depend on this document; this preregistration governs the Small-track and Main-track ESTIMATORS and their evaluation. Its purpose is to make the R-662B12 failure impossible to repeat: no rule is chosen after reading the labels, everything learnable is fit inside folds, and generalization is measured by holding out whole problems and whole models.

Data of record (committed, offline)

Permitted inference-time features (in-container: model_id string + problem texts only)

Computable from the 8B model on the given problem and self-generated, meaning-preserving perturbations of it:

  1. self-perturbation answer agreement: fraction of K self-generated semantic-preserving perturbations on which the model's final answer equals its answer on the original problem;
  2. self-consistency: agreement of the model's final answer across R repeated samples of the original problem;
  3. response uncertainty statistics of the original response (token-logprob margin, entropy, answer-token confidence), the cheap end used by the official uncertainty baseline.

The model_id string may enter ONLY as a fitted categorical prior (per-model intercept) with a mandatory out-of-model fallback, and only if it improves the binding held-out metric (below). It may never be a hard-coded name to label map.

Prohibited leakage (hard)

Fitting and evaluation

Controls (all preregistered)

Success and kill gates

What a positive result would and would not mean

A method that beats constants out-of-fold on public 8B data is evidence for a leaderboard submission, not a guarantee on the hidden set, because this hunt's own curation finding shows the public sample's class balance is not the hidden sample's. The report states this explicitly. Nothing here uses the reserved verification vocabulary; every claim is "measured on public data" and graded as such.

Amendment, 2026-08-22 (recorded after the original freeze, text above unchanged)

  1. Main-track binding split. The original text made leave-one-model-out binding on the assumption that the hidden set contains unseen models. The competition site ("The provided validation set covers all types of models contained in the test set") and the proposal paper (§1.6, "All model types included in the private test set will be covered in the public validation set") say otherwise. Leave-one-problem-out is therefore the binding split for the Main track; leave-one-model-out is still reported and is the prior's fallback case (unseen identifier -> non-robust). This is the split the original text already used for the Small track.
  2. Consequence for the permitted model_id prior. The categorical per-model prior was permitted above "only if it improves the binding held-out metric". Under leave-one-problem-out it does, on every public set (26/28, 39/41, 43/56 vs 19/28, 32/41, 32/56), so it ships as the Main-track entry. Its table is generated by protocol_reconstruction.py from the official val-sample labels, not written by hand, and the fallback is the better constant.
  3. Better constant. The original text chose the constant on the natural public distribution. The evaluation distribution is the clear-cut protocol (proposal §1.4), reconstructed in protocol_reconstruction.py; on it and on every organizer-released labelled set the better constant is always non-robust. The always-robust bundles are superseded, not deleted.
  4. Free gate, re-run on more data. The organizers' MATH releases (augmented-sample-math*, public, organizer-made, so not "extra labeled data" under the rules) give the 8B a minority class the AIMO sample lacks. The gate still fails: no single cheap perturbation type predicts failure under another out of fold. The Small-track GPU kill stands; no second attempt.