LAB-1 — a broken import in a hypothesis-lab test fail-closes all three trading sleeves, and the marker convention that is supposed to prevent exactly this cannot
status: open · raised panel run 29 (2026-09-28; change-reviewer catch on the 2026-09-27 lab commits — a defect row, not a research or capture proposal, so it does NOT discharge the scout's 8-week cadence clock; same treatment as BTB-1 and CQA-3) · class: trading safety (G2) + isolation hygiene (G5) · effort ~0.25 pd · horizon: live today, and the exposure grows with every lab edit · queue was at the 12-file cap when this was raised; see "The slot this costs"
Plain-language summary for an owner reading one paragraph. On 2026-09-27
you merged the hypothesis lab — 5,384 lines — straight to main. It was
written carefully: its tests are auto-tagged research so that a lab bug can
never stop the three sleeves from trading, which is the lesson of 2026-07-17,
when a rotted reporting test blocked momentum and meanrev for two days. That
tag does not do what it is written to do. The tag is applied while pytest is
collecting tests; if a lab test file fails to import at all, collection
dies before any tag exists, the pre-trade gate exits non-zero, and not one
trading-critical test ever runs. All three sleeves fail closed, for a
research tool's defect. Nothing is broken today. What is missing is the
mechanism that would keep it that way — and the one test that currently keeps
today safe has not executed since the 2026-09-19 billing outage.
Mechanism — measured at HEAD 864da4e, reproduced from scratch
Leg 1 — marker deselection happens AFTER collection
The three pre-trade gates are identical
(.github/workflows/paper-trading.yml:89, paper-trading-meanrev.yml:84,
paper-trading-vrp.yml:70):
python -m pytest tests/ -q -m "not observability and not research" --tb=line
tests/test_lab/conftest.py:23-29 applies the research marker in
pytest_collection_modifyitems — a hook that runs after every test module
has been imported. A module-level ImportError therefore never reaches the
marker at all.
Reproduced in a scratch tree built from nothing (not a copy of the repo), with one trading-critical test and one auto-marked lab test carrying a bad import:
$ python -m pytest tests/ -q -m "not observability and not research" --tb=no
ERROR tests/test_lab/test_lab_thing.py
!!!!!!!!!! Interrupted: 1 error during collection !!!!!!!!!!
EXIT=2
test_trading_critical_invariant did not run. The gate reds, and on a
trading day that is a sleeve that does not trade.
Leg 2 — the isolation claim is false at the import-graph level
src/thales/cli.py:5535 imports the lab sub-app at module scope:
from thales.lab.cli import lab_app # noqa: E402 — light module; heavy imports are per-command
Measured:
$ python -c "import sys, thales.cli; print(sorted(m for m in sys.modules if m.startswith('thales.lab')))"
['thales.lab', 'thales.lab.cli']
So thales run, thales halt and thales portfolio all load the lab. Two
places in the tree assert the opposite:
| file:line | text |
|---|---|
CLAUDE.md:148 | "HYPOTHESIS LAB (research only; no trading path imports it)" |
src/thales/lab/__init__.py:37 | "Nothing on any trading path imports this package." |
(The RESEARCH.md banner for the lab says the same.) A module-level failure
anywhere under thales.lab takes the whole CLI down — including thales halt,
the manual kill switch.
Why "nothing is broken today" is not reassurance
Today's exposure is nil because the lab imports only declared dependencies
(numpy, pandas, scipy, pyyaml, typer + stdlib). The property that keeps it
nil is enforced by tests/test_declared_dependencies.py, which is marked
observability and therefore runs only in fleet-digest.yml — which has
not executed since Actions was refused on 2026-09-19. So a 5,384-line package
landed on main with zero CI, is under active edit (both committed study runs
are stamped +lab-dirty, i.e. run from an uncommitted tree), and the single
test that would catch a broken or undeclared import in it is not currently
running.
This is the 2026-07-29 matplotlib shape — a package that arrived transitively, vanished silently, and killed the digest for three trading days — pointed at the trading gate instead of the digest.
Data plan
None. No data, no PiT question, no registry row, no trial. Two workflow lines and either a two-line source change or two doc corrections.
Test design
Both legs get a pin with a real negative control; neither is a doc edit alone, because a corrected sentence with nothing enforcing it re-creates the same undefended assertion.
Leg 1. Add --ignore=tests/test_lab to the three pre-trade gates
alongside the marker (belt and braces — -m cannot cover collection
errors). Pin it with a test that asserts all three gate commands carry both
the marker expression and the ignore.
Falsification, measured in the same scratch tree:
| gate command | trading test | lab breakage still loud? |
|---|---|---|
-m "not observability and not research" | never runs, EXIT=2 | — |
… --ignore=tests/test_lab | 1 passed, EXIT=0 | — |
-m "observability or research" (fleet-digest) | — | EXIT=2, still reds |
The third row is the control that matters: the fix must not buy trading-gate safety by making a lab breakage silent. It does not.
Leg 2. Either (a) make the sub-app mount lazy or wrap it so an unavailable
lab cannot kill the CLI, and add an import-graph test asserting thales.lab
is not in sys.modules after import thales.cli; or (b) if the eager
mount is preferred, correct both sentences and pin the actual invariant —
that thales.lab.cli imports nothing beyond stdlib + typer at module scope.
Pick one and pin it; do not just fix the prose.
Pre-registered kill criterion
This row dies if either of these turns out true:
- The hazard is unreachable. If a test is added showing that a
module-level
ImportErrorundertests/test_lab/cannot interrupt the pre-trade gate as it is actually invoked in the workflows — i.e. my reproduction does not transfer to the real gate — the row is wrong and closes. (I reproduced the mechanism in a scratch tree, not against the live workflow runner; that is the gap a falsifier should aim at.) - The lab is removed from
tests/or from the CLI. Iftests/test_lab/moves outside thetests/tree the gate walks, or the lab sub-app stops being mounted on the production CLI, both legs evaporate and the row closes as overtaken.
It does not die merely because the suite is green — "green today" is the state the row is about.
Effort and horizon
~0.25 pd, implementer-buildable (routine PRs touch workflows — #185/#186 do). Decidable immediately. The first execution of the changed pre-trade gate will be the first trading day after Actions resumes, expected 2026-10-01 — momentum's monthly alpha-selection day. Both suites are green at HEAD locally (pre-trade 1229 passed / 1 skipped; research tier 144 passed), so this row is not a blocker for 10-01; it is the guard that should exist before the next lab edit, not before the next trading day.
The slot this costs
research/queue/open/ held exactly 12 files — the pre-registered cap —
when this was raised, and research/queue/approved/ has never held an item in
its existence. So this row cannot be free: it evicts an incumbent, and it
lands in a queue nothing is being pulled from.
It was opened anyway because it is the only candidate this run that is live, unguarded, buildable, and backed by a realised precedent already paid for in this shop's own ledger. The adversarial reviewer nominated RWG-1 for eviction (the queue's oldest, raised run 11 on 2026-08-20; its leg 1 was already discharged by the owner that same day, so it is partially resolved, and its remaining leg is a guard corner case with no date and no active bleed). Alternate nominee CQA-3 — but evicting that one immediately fires CQA-2's reopen leg (c), a downstream cost RWG-1 does not carry.
The ranking call is triage's mechanical one on Monday, not the panel's. A nominee is named so the trade is explicit rather than deferred; the move is not made here.
Re-verified 2026-10-01 against main at 3d4f432 (appended, body above unchanged)
This PR was stranded by the outage and sat unmerged for three days while the
lab kept growing (e53fa72, a9cffcd, ada5bff — a 1926+ panel, three new
families, five confirmatory studies). Every claim above still holds, and the
exposure is larger, not smaller:
- Both legs intact. All three pre-trade gates still carry
-m "not observability and not research"with no--ignore(grep -c "ignore=tests/test_lab"→ 0).src/thales/cli.py:5535still importslab_appat module scope. Both isolation sentences are still present and still false. tests/test_lab/grew from 11 files to 13 (test_confirmatory.py,test_longrun.pyadded), so there is more surface that can fail to import inside the trading gate than when this row was written.- Line references that drifted (the substance did not): the meanrev gate
is now
:125(was:84), the vrp gate:110(was:70), and CLAUDE.md's isolation claim:149(was:148).cli.py:5535andlab/__init__.py:37are unchanged.
Nothing here changes the ask, the test design or the kill criterion.