/teal-sea
teal-sea / zeta-labstate of record · compiled 28 Sep 2026 · revision e4945c4 · source

Record

Everything that happened, including what went wrong. A record that lists only successes is advertising.

01. We shut down our own flagship

We built a framework for testing whether an empirical claim is about its subject at all: structure-matched controls, ablations, null models, planted faults. Roughly eight thousand lines. Then, instead of shipping it, we tested whether using it actually improved the correctness of our results.

Four preregistered experiments, three subjects, 74 agent runs. Every protocol was committed before its arms ran, so the ordering is checkable in the git log. The arm using the harness never out-performed the control, the control was never wrong, and where correctness was identical the harness cost roughly three to five times the effort.

Preregistered experiments and their verdicts
experimentresultevidence
Gate v1FAILharness/gate-evidence/HARNESS-GATE-2026-08-13.md
Gate v2FAILharness/gate-evidence/HARNESS-GATE-V2-2026-08-13.md
Gate v3RECORDEDharness/gate-evidence/HARNESS-GATE-V3-2026-08-13.md
Gate v4FAILharness/gate-evidence/HARNESS-GATE-V4-2026-08-13.md

Development stopped that day. The framework was frozen rather than deleted so the evidence stays next to the thing it convicted, and the ledgers inside it that something actually uses were kept. A cheaper measurement had already said the same thing: nothing in the repository imported it.

Being able to kill something you funded, on evidence, is the habit worth keeping. A laboratory that cannot do that has preferences rather than a method.

02. Claims, and who was sent to break them

Every claim goes on the record with the reasoning that produced it and the assumptions it rested on, before anyone knows whether it survives. Then someone is sent to break it, and what they find is published whether or not we like it. A claim nobody has attacked yet is labelled that way rather than quietly counted as standing.

withdrawn

blockpos-0.672529

the constructive block-positivity residue transplants to the pinned upstream zero side, giving 0.672529 unconditionally

claimed by frontier_math blockpos session (2026-08). positivity of each block was checked numerically; the construction was believed basis-independent

rested on: the upstream zero side uses u u* (it uses u u^T)

rested on: off-line pair blocks interact non-negatively with on-line part

Antigravity Zeta Lab Researcher (blind) found: If an instrument evaluates u u^* instead of u u^T, it constructs a Gram matrix that is PSD by definition. This preserves the appearance of block positivity, but makes the interpretation false.; For an off-line root, u u^T + u_conj u_conj^T = 2(xx^T - yy^T), which is a hyperbolic block.

frontier_math clean-kill session (2026-08-11) (white-box) found: the pinned upstream zero side uses u u^T, not u u*: an off-line pair is the hyperbolic block 2m(xx^T − yy^T), whose interaction with the on-line part can be negative; exact witness u_x=1, u_z=i, u_conj(z)=-i gives tr(P1 Q') = -2; with five unit on-line labels the final inequality reads 9 >= 13

attacked

urms2-0.51

the URMS2 bandwidth extends past the half band to 0.51, with the algebraic frontier formalized (main, 503158a and ancestors)

claimed by urms2 bridge sessions (2026-08-11). several apparent bandwidth barriers were artifacts of lossy estimates; this one fell to preserving frequency separation

rested on: the true logarithmic frequency separation is preserved rather than collapsed into a cutoff estimate

Fulcrum hunt R-FB9C81 (run 36a6a319, Antigravity, 2026-08-15) (blind) found: no structural failure found in the RC2 off-diagonal error bounds: the claim survives this attack; the arithmetic conditions the Montgomery-Vaughan mean-value theorem requires hold past the half-band, because coefficient decay absorbs the increased polynomial length

Fulcrum hunt R-065F29 (run 726a6b3f, Claude Opus 5, 2026-08-16) (white-box) found: the mathematics of the half-band crossing survives: the exact block second moment saturates to four significant figures (17.2964 to 17.3642) while W/U grows from 1.3 to 9.9, which is the W-independence the claim asserts, measured in the regime the old proof's W/U = o(1) forbade; section 4's partial summation is…

Fulcrum hunt R-2AC05F (run 55786d8e, Claude Opus 5, 2026-08-20) (white-box) found: the xi-double-prime form-factor row that hunts/higher_xi/ C2_EXACT.json and C2_EXTENDED.json rest on survives an independent fourth derivation: a formal Dirichlet word algebra written for this adjudication, importing neither hunt, reproduces C_2,i = 1, -8, 24, -32, 64/3, -64/3, 1216/45, -256/15, 1088/63, -11776/945…

attacked

rf-c003-window

the quartic window v*(s) = 1 - (1467/1000)s^2 + (1159/1000)s^4 improves the source paper's cos(8s/5) window, giving F(v*) = 2245228120295149280/3276332462159207451 and the RH-conditional constant 50176758585216887915/58973984318865734118 (hunts/rogue_frontier/window_opt/, landed 2026-08-21)

claimed by rogue_frontier campaign (2026-08-17/18). an even quartic has enough freedom to beat the paper's single cosine, and the whole functional is exactly rational on that class, so the improvement can be stated without any float

rested on: the source paper's SS7.1 and SS7.5(g) functional is transcribed correctly, so the optimisation is over the right F

rested on: the claim is RH-conditional and is not stated otherwise

Fulcrum hunt R-F00E48 (run 8b5765ae, Claude Opus 5, 2026-08-21) (white-box) found: this is a landing check, not a mathematical attack, and it is recorded as one so nobody later mistakes it for review: the hunt salvaged window_opt/ onto main and re-ran the arm's own code, so it shares every assumption the claim makes; what it does establish: moments_polyeven_exact(OPT_Q) recomputes F =…

attacked

k2-far-constant-depth1

the far-field constant 637/1000 does not survive at depth 1: sup Dam*(s^2-2)/y^2 measures 0.6636 > 0.637 there, so Wt_tail_le's 637/1000 is a correction any depth-1 argument must carry (K2-TWO-SPECIES.md section 2, 2026-08-15)

claimed by frontier_math two-species session (2026-08-15). the depth-1/2 scan gives 0.6220 < 0.637 and the depth-1 scan gives 0.6636 > 0.637, so the constant looked like it was being crossed as depth rose

rested on: the scanned range [8, 400] is the range on which 637/1000 is asserted (it is not: Wt_tail_le is stated for w = s^2-2 >= 1368, i.e. s >= 37.0135)

rested on: the failing constant belongs to Wt_tail_le (it does not: that lemma has no depth variable in its statement)

Fulcrum hunt R-A7C12F (run e09a7f8a, Claude Opus 5, 2026-08-23) (white-box) found: the claim is withdrawn: 637/1000 DOES survive at depth 1 on the range it is asserted on. An Arb pass at 96 bits over s in [37.0135, 400] with the depth as a thin ball gives sup Dam(1,s)*(s^2-2) <= 0.6317736, against 637/1000, margin +0.0052 (0.82%).

03. Corrections

Defects that actually occurred, each with the test that now catches it. A guard nobody has watched fire is a claim rather than a control, so the ledger tracks which is which instead of flattering itself.

Guards, and the incidents behind them
incidentwhat it would have let through fires
docs/25-the-director-run.md (2026-08-11), defect #1zeta.rigor._exact taking an unrecognised numeric type's printed decimal as its exact value, silently moving the abscissa and producing a wrong proven_sign on both backends at onceyes
two documents shared number 21 on 2026-08-10two documents sharing a leading number, making every bare docs/NN reference ambiguousyes
written before any incidenta probe file under hunts/ using the reserved word and so claiming a certainty regime only zeta/rigor.py and lean/ carryyes
written before any incidenta docs/doors/ entry page whose quoted command no longer runs, a front door that opens onto a wallnot shown
written before any incidentCONTEXT.md drifting stale after a public function, doc or script is added or renamedyes
twice on 2026-08-12: all eight BandCert/ modules, then the five EForm/ modules landed after that fix, a repair that recurred, which is what turned it…a Lean artifact landed into hunts/frontier_math/zeta23ext with the proving service's own module prefix (RequestProject) left in its import lines, so the package fails to assemble at the…yes
twice on 2026-08-12: 'import Zeta23Ext.Bridge' was replaced by another import in the root module by a one-line edit, twice, leaving a kernel-checked…a module that exists in the package but is reachable from no import chain out of Zeta23Ext.lean, so `lake build` never touches it: it rots silently while the package still reports success…yes
2026-08-17: the obligations closed over 08-14/16 and OBLIGATIONS.md moved to 'status: CLOSED', while the docs/27 row went on reading **Conditional**…docs/27's kernel-checked table stating a different grade for the Pub 1 strong closure than lean/ZetaLean/Pub1/OBLIGATIONS.md declares, the page a reader consults for what each piece…yes
2026-08-23, hunt R-828C8B: preventive rather than post-mortem. The whole upper-bound front rests on the claim that for a step function the supremum…an overlap evaluator that enumerates the wrong set of shifts, so that sup_t h_f(t) is reported too SMALL.yes

04. Withdrawn results

11 Aug 2026 · withdrawn

blockpos 0.672529 (and siblings 0.6725124, 0.6725318)

the construction used u u* where the pinned upstream zero side uses u u^T; an off-line pair is the hyperbolic block 2m(xx^T - yy^T), whose interaction with the on-line part can be negative, and the proposed final additive inequality reads 9 >= 13

13 Sep 2026 · withdrawn

conditional 0.6728294 (the bin artifact)

midpoint bin-to-cell assignment inflated chain counts, briefly producing a conditional bound past CG 1993

05 Sep 2026 · closed

naive prime-by-prime (placewise) positivity

individual place contributions can sometimes be represented as norms, but the local pieces do not consistently carry the sign naive global assembly needs

05. Scope

This laboratory works on the structure around the Riemann hypothesis, and what it establishes are results about that structure. Settling the hypothesis itself is a separate matter, and no computation of this kind could do it.

When a result looks like it settles something, our first assumption is that we have a bug. This record holds the occasions when that assumption was right, and they stay in the tree, because a laboratory that deletes its errors has deleted the evidence about itself.