2026-09-14 — the dark-routine exception could never fire
Daily audit, 2026-09-14. Observability finding; no trading logic touched.
What happened
On 2026-09-12 the fleet digest moved from a daily email to weekly (Friday) plus any exception day (#163). The design was explicit that this must not cost detection latency, and said so in the code itself:
cadence is reduced for QUIET days only: any day carrying an exception still mails the same evening. Fewer emails, same loudness — never the trade the request literally asked for, and this comment is why. —
fleet_digest.py, abovedigest_exceptions
The docstring names the standing example: "A dark routine is the standing example: nothing is red, and nobody is watching."
2026-09-14 was the first weekday the policy was live. The digest ran (run 34889997940, 19:58Z), computed routine liveness correctly, and printed:
⚠ 2 routine(s) DARK — scheduled runs produced no beacon:
daily audit: 5 missed since 9/07 (last ran 9/04)
research implementer: 1 missed since 9/09 (last ran 9/02)
email withheld by cadence policy — quiet day off-cadence.
It rendered the alarm and then withheld the email carrying it — on the exact failure the exception path was written for, while the platform's own safety retrospective had been dark for five scheduled fires.
Root cause — one key, in the wrong place
_gather_pipeline_status sets out["routines"] (fleet_digest.py:423), and
gather_fleet_digest returns that dict nested under "pipeline". Both
renderers read it there (pl.get("routines"), :884 and :979).
digest_exceptions read f.get("routines") — a top-level key the gather
never writes. It therefore always saw []. The dark-routine branch had never
been able to fire in production, on any day, since it was written.
The other three exception sources read correct paths (f["fleet"]["red"],
f["fleet"]["no_run_sleeves"], f["sleeves"][…]["data"]["failed_orders"]),
so this was the only dead branch.
Proven by running the shipped function against the shape the 09-14 digest produced, with the top-level shape as the control:
digest_exceptions({"pipeline": {"routines": [ …5 missed… ]}}) -> []
should_email(…, "weekly", 5, 2026-09-14) -> (False, 'quiet day off-cadence')
# control — same rows moved to the top level, where the code looked:
digest_exceptions({"routines": [ …5 missed… ]})
-> ['2 routine(s) dark (daily audit, research implementer)']
Why the test suite did not catch it
tests/test_execution/test_digest_cadence.py did pin this behaviour — the
file exists for exactly this worry, and its module docstring names the
five-weekday darkness as its motivation. But its fixtures built routines at
the top level, matching the buggy reader rather than the gather. The test and
the defect agreed with each other, so the suite was green while the feature
was dead.
The lesson worth keeping: a fixture that does not match what the gather actually returns does not pin behaviour — it certifies the bug. This is the markdown-is-executable presumption applied to test data: a fixture is an assertion about production shape, and it deserves a negative control like any other claim.
Fix
digest_exceptionsnow readsf["pipeline"]["routines"], the same path both renderers use.- Fixtures corrected to the gather's real shape.
- Four new tests: the 09-14 scenario end-to-end; a direction-pinning case asserting the top-level key is NOT what the decision reads (so the old fixture shape cannot quietly return); an invariant tying the rendered DARK line to the send decision so the two readers cannot drift again; and a negative control that routines reporting on schedule still withhold.
- Negative control run: with the one-line source fix reverted, 4 tests fail; with it, 14 pass and the execution suite is 577 passed.
Known residual (owner's call, not fixed here)
_gather_routine_liveness leaves missed=None when a beacon branch cannot be
read (e.g. the gh call fails). The renderers report those separately as
? N routine(s) with no readable beacon, but digest_exceptions does not
treat them as exceptions — so an unreadable beacon stays invisible for up to
seven days under weekly cadence. Making it an exception risks a daily email
whenever gh is flaky, so the trade-off is the owner's. Current behaviour is
pinned by a test so any change is deliberate.
Impact while it was live
2026-09-12 → 2026-09-14, one weekday of exposure (09-14). The cost was one withheld exception email — but it was withheld on a day that genuinely qualified, which is the whole failure mode the feature was built to prevent. Detection latency for a dark routine was, in effect, still up to seven days.
Pre-merge review follow-up (append only; nothing above rewritten)
Research panel run 23 (2026-09-15) accepted the fix and found the guard
short of what 4b30d99 (#169) asked of DGX-1 step 3: every new fixture was
still a hand-built literal, so the READER key was pinned and the WRITER key
nowhere. Re-measured before merge: renaming out["routines"] in
_gather_pipeline_status, or hoisting routines to the top level of the
gather's return, left all of tests/ green (1400 passed). The sentence above
calling test_a_rendered_DARK_line_and_the_email_decision_cannot_disagree an
invariant between the two readers overclaims: it hands rows straight to the
shared formatter, bypassing each renderer's own extraction.
Added before merge: two hermetic tests in test_digest_cadence.py that drive
the production join itself (generate_fleet_digest → gather_fleet_digest( include_pipeline=True) → should_email → send_email, stubbing only cwd,
gh, the beacon lookup and SMTP) — the 09-14 scenario must mail with the
DARK line in both bodies, and an all-on-schedule control must stay withheld
and still render its liveness line. The dark case goes red under the fix
revert, both producer mutations, and a renderer key rename.
Scope: this covers the routine-liveness leg only. It does not discharge DGX-1 step 3 (every renderer alarm line ↔ an exception), which stays open.