# LAB-1 — a broken import in a hypothesis-lab test fail-closes all three trading sleeves, and the marker convention that is supposed to prevent exactly this cannot

status: open · raised panel run 29 (2026-09-28; change-reviewer catch on the
2026-09-27 lab commits — a defect row, **not** a research or capture proposal,
so it does **NOT** discharge the scout's 8-week cadence clock; same treatment
as BTB-1 and CQA-3) · class: trading safety (G2) + isolation hygiene (G5) ·
effort ~0.25 pd · horizon: live today, and the exposure grows with every lab
edit · **queue was at the 12-file cap when this was raised; see "The slot this
costs"**

---

**Plain-language summary for an owner reading one paragraph.** On 2026-09-27
you merged the hypothesis lab — 5,384 lines — straight to `main`. It was
written carefully: its tests are auto-tagged `research` so that a lab bug can
never stop the three sleeves from trading, which is the lesson of 2026-07-17,
when a rotted reporting test blocked momentum and meanrev for two days. **That
tag does not do what it is written to do.** The tag is applied while pytest is
*collecting* tests; if a lab test file fails to *import* at all, collection
dies before any tag exists, the pre-trade gate exits non-zero, and **not one
trading-critical test ever runs**. All three sleeves fail closed, for a
research tool's defect. Nothing is broken today. What is missing is the
mechanism that would keep it that way — and the one test that currently keeps
today safe has not executed since the 2026-09-19 billing outage.

## Mechanism — measured at HEAD `864da4e`, reproduced from scratch

### Leg 1 — marker deselection happens AFTER collection

The three pre-trade gates are identical
(`.github/workflows/paper-trading.yml:89`, `paper-trading-meanrev.yml:84`,
`paper-trading-vrp.yml:70`):

```
python -m pytest tests/ -q -m "not observability and not research" --tb=line
```

`tests/test_lab/conftest.py:23-29` applies the `research` marker in
`pytest_collection_modifyitems` — a hook that runs *after* every test module
has been imported. A module-level `ImportError` therefore never reaches the
marker at all.

Reproduced in a scratch tree built from nothing (not a copy of the repo), with
one trading-critical test and one auto-marked lab test carrying a bad import:

```
$ python -m pytest tests/ -q -m "not observability and not research" --tb=no
ERROR tests/test_lab/test_lab_thing.py
!!!!!!!!!! Interrupted: 1 error during collection !!!!!!!!!!
EXIT=2
```

`test_trading_critical_invariant` **did not run**. The gate reds, and on a
trading day that is a sleeve that does not trade.

### Leg 2 — the isolation claim is false at the import-graph level

`src/thales/cli.py:5535` imports the lab sub-app at module scope:

```python
from thales.lab.cli import lab_app  # noqa: E402 — light module; heavy imports are per-command
```

Measured:

```
$ python -c "import sys, thales.cli; print(sorted(m for m in sys.modules if m.startswith('thales.lab')))"
['thales.lab', 'thales.lab.cli']
```

So `thales run`, `thales halt` and `thales portfolio` all load the lab. Two
places in the tree assert the opposite:

| file:line | text |
|---|---|
| `CLAUDE.md:148` | "HYPOTHESIS LAB (research only; **no trading path imports it**)" |
| `src/thales/lab/__init__.py:37` | "**Nothing on any trading path imports this package.**" |

(The RESEARCH.md banner for the lab says the same.) A module-level failure
anywhere under `thales.lab` takes the whole CLI down — including `thales halt`,
the manual kill switch.

### Why "nothing is broken today" is not reassurance

Today's exposure is nil because the lab imports only declared dependencies
(`numpy, pandas, scipy, pyyaml, typer` + stdlib). The property that keeps it
nil is enforced by `tests/test_declared_dependencies.py`, which is marked
`observability` and therefore runs **only in `fleet-digest.yml`** — which has
not executed since Actions was refused on 2026-09-19. So a 5,384-line package
landed on `main` with zero CI, is under active edit (both committed study runs
are stamped `+lab-dirty`, i.e. run from an uncommitted tree), and the single
test that would catch a broken or undeclared import in it is not currently
running.

This is the 2026-07-29 matplotlib shape — a package that arrived transitively,
vanished silently, and killed the digest for three trading days — pointed at
the trading gate instead of the digest.

## Data plan

None. No data, no PiT question, no registry row, no trial. Two workflow lines
and either a two-line source change or two doc corrections.

## Test design

Both legs get a pin with a real negative control; neither is a doc edit alone,
because a corrected sentence with nothing enforcing it re-creates the same
undefended assertion.

**Leg 1.** Add `--ignore=tests/test_lab` to the three pre-trade gates
*alongside* the marker (belt and braces — `-m` cannot cover collection
errors). Pin it with a test that asserts all three gate commands carry both
the marker expression and the ignore.

Falsification, measured in the same scratch tree:

| gate command | trading test | lab breakage still loud? |
|---|---|---|
| `-m "not observability and not research"` | **never runs**, EXIT=2 | — |
| `… --ignore=tests/test_lab` | **1 passed**, EXIT=0 | — |
| `-m "observability or research"` (fleet-digest) | — | **EXIT=2, still reds** |

The third row is the control that matters: the fix must not buy trading-gate
safety by making a lab breakage silent. It does not.

**Leg 2.** Either (a) make the sub-app mount lazy or wrap it so an unavailable
lab cannot kill the CLI, and add an import-graph test asserting `thales.lab`
is **not** in `sys.modules` after `import thales.cli`; or (b) if the eager
mount is preferred, correct both sentences and pin the *actual* invariant —
that `thales.lab.cli` imports nothing beyond stdlib + typer at module scope.
Pick one and pin it; do not just fix the prose.

## Pre-registered kill criterion

This row dies if **either** of these turns out true:

1. **The hazard is unreachable.** If a test is added showing that a
   module-level `ImportError` under `tests/test_lab/` *cannot* interrupt the
   pre-trade gate as it is actually invoked in the workflows — i.e. my
   reproduction does not transfer to the real gate — the row is wrong and
   closes. (I reproduced the mechanism in a scratch tree, not against the live
   workflow runner; that is the gap a falsifier should aim at.)
2. **The lab is removed from `tests/` or from the CLI.** If `tests/test_lab/`
   moves outside the `tests/` tree the gate walks, or the lab sub-app stops
   being mounted on the production CLI, both legs evaporate and the row closes
   as overtaken.

It does **not** die merely because the suite is green — "green today" is the
state the row is about.

## Effort and horizon

~0.25 pd, implementer-buildable (routine PRs touch workflows — #185/#186 do).
Decidable immediately. The first execution of the changed pre-trade gate will
be the first trading day after Actions resumes, expected **2026-10-01 —
momentum's monthly alpha-selection day**. Both suites are green at HEAD
locally (pre-trade 1229 passed / 1 skipped; research tier 144 passed), so
**this row is not a blocker for 10-01**; it is the guard that should exist
before the next lab edit, not before the next trading day.

## The slot this costs

`research/queue/open/` held exactly **12** files — the pre-registered cap —
when this was raised, and `research/queue/approved/` has never held an item in
its existence. So this row cannot be free: it evicts an incumbent, and it
lands in a queue nothing is being pulled from.

It was opened anyway because it is the only candidate this run that is live,
unguarded, buildable, and backed by a realised precedent already paid for in
this shop's own ledger. The adversarial reviewer nominated **RWG-1** for
eviction (the queue's oldest, raised run 11 on 2026-08-20; its leg 1 was
already discharged by the owner that same day, so it is partially resolved,
and its remaining leg is a guard corner case with no date and no active
bleed). Alternate nominee **CQA-3** — but evicting that one immediately fires
CQA-2's reopen leg (c), a downstream cost RWG-1 does not carry.

**The ranking call is triage's mechanical one on Monday, not the panel's.** A
nominee is named so the trade is explicit rather than deferred; the move is
not made here.

---

## Re-verified 2026-10-01 against `main` at `3d4f432` (appended, body above unchanged)

This PR was stranded by the outage and sat unmerged for three days while the
lab kept growing (`e53fa72`, `a9cffcd`, `ada5bff` — a 1926+ panel, three new
families, five confirmatory studies). Every claim above still holds, and the
exposure is **larger**, not smaller:

- **Both legs intact.** All three pre-trade gates still carry
  `-m "not observability and not research"` with no `--ignore`
  (`grep -c "ignore=tests/test_lab"` → **0**). `src/thales/cli.py:5535` still
  imports `lab_app` at module scope. Both isolation sentences are still
  present and still false.
- **`tests/test_lab/` grew from 11 files to 13** (`test_confirmatory.py`,
  `test_longrun.py` added), so there is more surface that can fail to import
  inside the trading gate than when this row was written.
- **Line references that drifted** (the substance did not): the meanrev gate
  is now `:125` (was `:84`), the vrp gate `:110` (was `:70`), and CLAUDE.md's
  isolation claim `:149` (was `:148`). `cli.py:5535` and
  `lab/__init__.py:37` are unchanged.

Nothing here changes the ask, the test design or the kill criterion.
