LAB-3 (folded with LAB-4) — the lab's committed study YAML is not pinned to the ledger's recorded spec, and the vault's alpha ladder is not drawn from the declared budget
status: dismissed at panel run 30 (2026-09-29), raised the same run by the
change-reviewer lens on the three un-CI'd lab commits (e53fa72, a9cffcd,
ada5bff, 2026-09-29) · class: registration hygiene (G5) · the defects are
real and reproduced; the dismissal is about exposure and queue cost, not
about correctness
Plain-language summary. Two genuine gaps were found in the new hypothesis lab, and both were reproduced. (1) The study specification files that record what each experiment promised in advance are not checked against the copy stored in the lab's permanent ledger — so one could be edited after the verdict was recorded and every test would still pass. (2) The "vault" (the one-shot holdout test) computes its pass bar from a default in the config rather than from the error budget the shop declared, and a study can set that number for itself. Both were dismissed anyway, because the automatic-merge guard refuses to merge any change to the files at risk — so no unattended routine can reach them — and because nothing currently rides on the records in question.
What was measured (so nobody re-derives it)
LAB-3 — the YAML is not pinned to the ledger. The ledger records both
spec_text and spec_sha256 for every run (src/thales/lab/confirmatory.py:316-318,
src/thales/lab/study.py:538), but nothing compares either to the file on
disk. tests/test_lab/test_repo_state.py:57-61 only checks that each committed
YAML still expands; :64-75 re-derives rule ids from the ledger's own
spec_text, never from the file. results/lab/* is gitignored except
ledger.jsonl.
Reproduced twice (change-reviewer lens, then the adversarial reviewer, both in
scratch copies): retro-editing research/lab/studies/turn_of_month_1927.yaml
after its verdict was charged — loosening min_same_sign 2→1,
stress_cost_bps 10→1, cost_bps 5→0 — leaves pytest tests/test_lab -q at
194 passed.
The fix, if it is ever wanted, is ~8 lines and is currently green with zero
false positives: at HEAD all 8 ledger run rows' spec_sha256 match
sha256(file)[:16]; after the retro-edit turn_of_month_1927 reads ledger
f9f12d9c4001bfbb vs file e044a65eb7e1ff25. Note for whoever writes it:
spec_sha256 is a 16-hex truncation, not a full digest — a naive
full-length comparison fails on every row.
LAB-4 — the vault's alpha is not the declared budget's.
config/lab.yaml:43-63 declares total 0.05 = exploratory 0.025 + confirmatory
0.025, and check_budget (study.py:78-92) refuses a config that disagrees;
run_study:463 overrides a study's own value. The vault does not:
src/thales/lab/cli.py:585-587 computes fwer = spec.stats.get("fwer_alpha") or 0.05
then alpha = fwer / 2**k, keyed on defaults.fwer_alpha rather than the
ledger's budget row. That series sums to 0.05 — the entire declared total — on
top of the 0.025 + 0.025 already allocated. And fwer_alpha is in
spec.py:81 _STAT_KEYS with no bound, so a spec declaring fwer_alpha: 0.5
parses clean and gives a first-opening bar of p ≤ 0.25 — which the confirmatory
path explicitly forbids (confirmatory.py:173-175, "a study does not choose
its own bar").
Second leg: CLAUDE.md:422 and research/lab/README.md:56 both promise "an
owner-approved, dated budget amendment", but no amendment path exists —
check_budget:85-91 refuses a differing config, cli.py:437-440 refuses a
second --declare, ledger.budget() returns rows[0], and
test_repo_state.py:91 asserts "one budget, ever". The documented door is
welded shut.
Why dismissed — the measurement that decided it
The adversarial reviewer ran the auto-merge guard against the artifacts at risk, which nobody else had done:
research/lab/studies/turn_of_month_1927.yaml -> DENY: not a record file
results/lab/ledger.jsonl -> DENY: not a record file
research/lab/README.md -> ALLOW
No routine can land a retro-edit to a study YAML or to the ledger unattended. The threat reduces to a human-authored, human-merged change. That is the whole difference from FRZ-1 (opened the same run), where the guard says ALLOW and a routine has already used the path.
Three further facts make the exposure nil today:
- All five confirmatory verdicts are KILL or INCONCLUSIVE. Nothing is promoted on them, so loosening a bar retroactively would change no decision.
- The reserve is exhausted, 5 of 5. The population of future pre-registrations that could be loosened is zero until you act.
- Vault openings to date: 0 (
n_confirmations= 0 inresults/lab/ledger.jsonl), so LAB-4's alpha ladder has never fired.
Against that: research/queue/open/ sits at its 12-item cap, the owner
approval gate fired its 8-of-8 never-used kill counter on 2026-09-28, and
triage has reported DRAIN ZERO for three consecutive weeks. A guard on a
channel no routine can reach, protecting records nothing rides on, does not
outrank an incumbent row.
Reopen conditions — any one of these
Each is an event, not a judgement call, and each marks the moment the fix is both cheapest and actually necessary:
- (a) A sixth confirmatory study, or any budget amendment, is authorized. The fix must land before it runs — this is the one condition with real urgency, because a new pre-registration is exactly what the pin protects.
- (b)
ALLOWED_PATTERNSinscripts/automerge_guard.pyis widened to admitresearch/lab/studies/*.yamlorresults/lab/ledger.jsonl. That single edit converts this from human-merged to unattended and reverses the dismissal. - (c) Any lab rule enters the house promotion road (pre-registered battery → sleeve spec → forward paper clock). Then a verdict starts carrying money and the record must be tamper-evident.
- (d)
thales lab confirmis run for the first time — i.e.ledger.confirmationsbecomes non-empty — which arms LAB-4's alpha ladder for real.
Also ruled this run, recorded here so they are not re-derived
- The "budget amendment" wording is currently a promise with no mechanism.
Whichever way (a) is answered, either build
lab budget --amend(a dated superseding row thatledger.budget()reads last) or strike the phrase fromCLAUDE.md:422andresearch/lab/README.md:56. Leaving a documented door that is welded shut is its own small rot. stats.fwer_alphabeing unbounded at parse is the cheapest half of LAB-4 (one bound,≤ defaults.fwer_alpha) and could ride any later lab PR without waiting for a reopen.