The escalation channel had no retry — found on day 5 of the billing outage
Daily audit, 2026-09-25. Observability finding; no trading logic touched.
What was observed
GitHub Actions has refused every job in this repository since 2026-09-21 ("recent account payments have failed or your spending limit needs to be increased" — re-read verbatim today off the 18:49Z momentum run). Five trading days lost. The audit reported this on days 1 and 4 and escalated it both times.
The escalations never arrived. Not because the audit failed to write them — both
reports sit on outbox/daily-audit, and a third was written today — but because
the thing that turns them into email is itself a GitHub Actions job.
What it turned out to be
Three properties of routine-outbox.yml, each reasonable alone, combining into
an unrecoverable silent failure:
- It triggers only on
pushtooutbox/**. Restoring billing does not re-fire a push that already happened. A refused send is never retried. - It mails
message.mdat HEAD. Every report overwrites that one file, so even a manual re-run would have mailed only the newest; the earlier reports' text — still present in git history — would never have been read. - Nothing compares report commits against the
sent-log.mdmarkers. The daily audit's check 11 does this by hand and found the divergence; it has no way to deliver what was missed.
So the channel that exists to say "the platform is down, here is the one thing only you can do" fails silently, and unrecoverably, in precisely the outage that makes it matter. The audit's own detector saw it on day 1 and the reports still went nowhere for five days.
This is the platform's signature failure shape — truth divergence between layers — turned on the reporting layer itself.
The second defect, found while fixing the first
A push event runs the workflow file from the pushed ref. The outbox/*
and beacon/* branches each carry a full repo tree snapshotted when they were
created and never refreshed: on 2026-09-25, 324 files behind main.
So a relay fix merged to main does not reach the relay at all until the branch
is merged forward. A routine that ships a relay fix and walks away has shipped
nothing — a fix-that-never-activates, the same class as the finding it fixes.
Verified benign for the present: the copy of routine-outbox.yml on
outbox/daily-audit is byte-identical to main's, so no past fix to this file
was silently lost. It was luck — that file had not been edited since the
branches were made.
What was done
scripts/outbox_backlog.py + tests/test_execution/test_outbox_backlog.py,
wired into routine-outbox.yml. The relay now mails the backlog rather than
HEAD:
report commits = commits touching message.md
sent = everything up to and including the newest sent-marker
BACKLOG = report commits AFTER the newest sent-marker, oldest first
Each is read at its own commit, so an overwritten message.md no longer
loses anything, and each carries a [DELAYED] subject tag and a
written-vs-mailed header — an undated stale report reads as today's state, which
is the very divergence this routine exists to catch. The first successful relay
run after any outage flushes everything that was missed.
Against the staleness trap, the decision rule is loaded from origin/main at
run time rather than from the checked-out branch, so the one-time branch refresh
buys every future change to the rule for free.
Verification (this platform's rule is that a detector is only as good as its negative controls):
- against the real branch histories, not fixtures: the model resolves to
exactly the two unsent reports on
outbox/daily-audit(after marker0b2dc8d) and exactly the one onoutbox/triage(after53a798f); - end-to-end with SMTP stubbed, on a worktree of the real branch: two emails,
correct headlines extracted past the relay's own prepended note, correct
[DELAYED]labels (4.0 days and 24.0 hours); - negative control, the one that matters: a marker-only push sends nothing and writes no marker commit — this is what stops the flush looping, since the marker push re-triggers the workflow;
- negative control: with the logic unavailable, the relay mails exactly one email (HEAD), which is the pre-change behaviour. Fail toward one email, never toward silence is the relay's founding rule and it still holds for every failure mode of the flush;
- a subject containing the phrase
outbox: sent-markeris still treated as a report — the marker test is anchored, not searched. Two past detector defects here were substring matches; - 1229 trading-critical tests and 181 observability tests pass. The new tests
are marked
observability, so a failure here can never fail-close trading (the standing rule after #176).
What remains
- The fix cannot activate while Actions is refused, and once the PR merges
the
outbox/*branches must be merged forward frommainonce for this file to take effect. Recorded inops/DAILY_AUDIT.mdcheck 11 so the next audit carries it rather than trusting memory. - The flush recovers reports; it cannot recover urgency. A five-day-late "fix the billing" is a record, not an alarm. During an Actions outage the only channel that reaches the operator is the routine's own push notification — which is how days 4 and 5 actually reached him. That residual is architectural and unchanged: no watcher exists outside GitHub Actions and the routine scheduler.