/teal-sea
teal-sea / zeta-labstate of record · compiled 28 Sep 2026 · revision e4945c4 · source

Library · hunts/r_365c6c/RESULTS.md

Hunt R-365C6C: file-type boundary of the hunt reserved-word guard

1,022 words · 72 lines · source

Status: settled. The guard tests/test_hunt_probe_discipline.py::test_no_hunt_claims_the_reserved_word enforces an explicit suffix whitelist: path.suffix.lower() in {".py", ".md", ".json"}. Every file type outside this three-extension set passes unconditionally, allowing an overclaim to pass undetected in .txt, .lean, .sh, .toml, .yaml, .csv, .rst, .tex, .html, .jsonl, or extensionless files.

Reproduce: python3 hunts/r_365c6c/probe.py (~40 s, standard library and pytest only). Data: results.json.

Controls

The sandbox harness in probe.py applies three strict controls:

  1. Empty specimen control: The sandbox guard is executed against an empty specimen directory and passes (exit 0). Every failure below is caused strictly by the planted specimen.
  2. Restore control: After each specimen run, the specimen is deleted and the sandbox is verified to return to clean green status (exit 0) before proceeding to the next specimen.
  3. Clean negative controls: Clean files with sanctioned vocabulary (clean.py, clean.md, clean.json, clean.txt, clean.lean) are verified to produce no false alarms (0/5 caught).

Result 1: Census of the hunts/ tree

An exhaustive census of all 447 files under hunts/ in the repository:

ExtensionCountPercentageGuard StatusNotes
.py16737.4%ScannedPython probe scripts and modules
.md13830.9%ScannedDocumentation, MISSION.md, RESULTS.md
.json6614.8%ScannedMeasurement data and HANDBACK files
Scanned Subtotal37183.0%CoveredChecked by lexical scan
.lean6815.2%UnscannedLean 4 formal proof modules (e.g. frontier_math)
.png40.9%UnscannedBinary plot images
(no ext)20.4%Unscanned.gitignore, lean-toolchain
.sh10.2%UnscannedShell build script (assemble.sh)
.toml10.2%UnscannedLean package config (lakefile.toml)
Unscanned Subtotal7617.0%BlindCompletely bypassed
Total447100.0%

Key findings from the corpus scan:

Result 2: The Specimen Battery

Running the unmodified guard against 40 planted specimens in the isolated sandbox:

CategoryTested (n)CaughtMissedDetection RateExamples
Scanned positive controls660100%.py, .md, .json, uppercase .PY, .MD, .JSON
Scanned clean controls3030% (clean)clean.py, clean.md, clean.json (no false alarms)
Unscanned clean controls2020% (clean)clean.txt, clean.lean (no false alarms)
Plain text & data formats4040%.txt, .log, .out, .dat (all missed)
Table data formats2020%.csv, .tsv (all missed)
Structured config & interchange8080%.yaml, .yml, .toml, .ini, .cfg, .xml, .jsonl, .ndjson (all missed)
Code & script formats100100%.lean, .sh, .bash, .zsh, .c, .cpp, .rs, .go, .js, .ts (all missed)
Document & markup formats6060%.tex, .bib, .rst, .html, .htm, .svg (all missed)
Extensionless files3030%NOTES, Makefile, Dockerfile (all missed)
Compound extensions2020%.json.bak, .py.tmp (all missed)
Dotfiles / hidden files2020%.notes.txt, .scratch (all missed)
Binary formats2020%.png, .npy (all missed)

What I chose and why

  1. Unmodified guard execution: Rather than simulating or re-implementing the guard logic, probe.py copies tests/test_hunt_probe_discipline.py into a sandbox tempdir and executes pytest against it for every specimen. The reported detection status is the test process exit code itself.
  2. Comprehensive format battery: Instead of testing .txt alone, the battery tests 35 non-scanned variants across text, configs, scripts, proofs, markup, extensionless files, compound extensions, and dotfiles.
  3. Non-modification of the guard test: Widening the guard's extension filter in tests/test_hunt_probe_discipline.py is outside the scoped write permission of this hunt. The boundary is measured, documented, and recorded in harness/departments/guard_ledger.py.

What could not be settled

Loose threads

  1. Lean proof files in hunts are completely unscanned. hunts/frontier_math/ contains 68 .lean files where proof claims and certificates are authored. Why it might matter: an overclaim in a .lean comment or theorem docstring inside a probe bypasses the guard entirely. First step: evaluate whether .lean should be added to the guard's whitelist in tests/test_hunt_probe_discipline.py.
  2. JSON Lines (.jsonl) logs are not scanned. .jsonl files are increasingly used for structured logs and telemetry, but path.suffix.lower() evaluates to ".jsonl", which does not match ".json". Why it might matter: probes recording claims in .jsonl telemetry logs are blind to the guard. First step: add ".jsonl" and ".ndjson" to the allowed extension set in test_no_hunt_claims_the_reserved_word.
  3. Compound extensions like .json.bak and .py.tmp slip through. Python's pathlib.Path.suffix only returns the final extension. Why it might matter: backup files or temporary scratch dumps containing overclaims are ignored. First step: consider checking all suffixes (path.suffixes) or checking whether any suffix matches the whitelist.