# The escalation channel had no retry — found on day 5 of the billing outage

*Daily audit, 2026-09-25. Observability finding; no trading logic touched.*

## What was observed

GitHub Actions has refused every job in this repository since 2026-09-21
("recent account payments have failed or your spending limit needs to be
increased" — re-read verbatim today off the 18:49Z momentum run). Five trading
days lost. The audit reported this on days 1 and 4 and escalated it both times.

The escalations never arrived. Not because the audit failed to write them — both
reports sit on `outbox/daily-audit`, and a third was written today — but because
the thing that turns them into email is itself a GitHub Actions job.

## What it turned out to be

Three properties of `routine-outbox.yml`, each reasonable alone, combining into
an unrecoverable silent failure:

1. **It triggers only on `push` to `outbox/**`.** Restoring billing does not
   re-fire a push that already happened. A refused send is never retried.
2. **It mails `message.md` at HEAD.** Every report overwrites that one file, so
   even a manual re-run would have mailed only the newest; the earlier reports'
   text — still present in git history — would never have been read.
3. **Nothing compares report commits against the `sent-log.md` markers.** The
   daily audit's check 11 does this by hand and found the divergence; it has no
   way to *deliver* what was missed.

So the channel that exists to say "the platform is down, here is the one thing
only you can do" fails silently, and unrecoverably, in precisely the outage that
makes it matter. The audit's own detector saw it on day 1 and the reports still
went nowhere for five days.

This is the platform's signature failure shape — truth divergence between
layers — turned on the reporting layer itself.

## The second defect, found while fixing the first

A `push` event runs the workflow file **from the pushed ref**. The `outbox/*`
and `beacon/*` branches each carry a full repo tree snapshotted when they were
created and never refreshed: on 2026-09-25, 324 files behind `main`.

So a relay fix merged to `main` does not reach the relay at all until the branch
is merged forward. A routine that ships a relay fix and walks away has shipped
nothing — a fix-that-never-activates, the same class as the finding it fixes.
Verified benign for the present: the copy of `routine-outbox.yml` on
`outbox/daily-audit` is byte-identical to `main`'s, so no past fix to this file
was silently lost. It was luck — that file had not been edited since the
branches were made.

## What was done

`scripts/outbox_backlog.py` + `tests/test_execution/test_outbox_backlog.py`,
wired into `routine-outbox.yml`. The relay now mails the **backlog** rather than
HEAD:

    report commits  = commits touching message.md
    sent            = everything up to and including the newest sent-marker
    BACKLOG         = report commits AFTER the newest sent-marker, oldest first

Each is read at **its own commit**, so an overwritten `message.md` no longer
loses anything, and each carries a `[DELAYED]` subject tag and a
written-vs-mailed header — an undated stale report reads as today's state, which
is the very divergence this routine exists to catch. The first successful relay
run after any outage flushes everything that was missed.

Against the staleness trap, the decision rule is loaded from `origin/main` at
run time rather than from the checked-out branch, so the one-time branch refresh
buys every future change to the rule for free.

**Verification** (this platform's rule is that a detector is only as good as its
negative controls):

- against the **real** branch histories, not fixtures: the model resolves to
  exactly the two unsent reports on `outbox/daily-audit` (after marker
  `0b2dc8d`) and exactly the one on `outbox/triage` (after `53a798f`);
- end-to-end with SMTP stubbed, on a worktree of the real branch: two emails,
  correct headlines extracted past the relay's own prepended note, correct
  `[DELAYED]` labels (4.0 days and 24.0 hours);
- negative control, the one that matters: a **marker-only push sends nothing**
  and writes no marker commit — this is what stops the flush looping, since the
  marker push re-triggers the workflow;
- negative control: with the logic unavailable, the relay mails exactly **one**
  email (HEAD), which is the pre-change behaviour. *Fail toward one email, never
  toward silence* is the relay's founding rule and it still holds for every
  failure mode of the flush;
- a subject containing the phrase `outbox: sent-marker` is still treated as a
  report — the marker test is anchored, not searched. Two past detector defects
  here were substring matches;
- 1229 trading-critical tests and 181 observability tests pass. The new tests
  are marked `observability`, so a failure here can never fail-close trading
  (the standing rule after #176).

## What remains

- **The fix cannot activate while Actions is refused**, and once the PR merges
  the `outbox/*` branches must be merged forward from `main` once for this file
  to take effect. Recorded in `ops/DAILY_AUDIT.md` check 11 so the next audit
  carries it rather than trusting memory.
- **The flush recovers reports; it cannot recover urgency.** A five-day-late
  "fix the billing" is a record, not an alarm. During an Actions outage the only
  channel that reaches the operator is the routine's own push notification —
  which is how days 4 and 5 actually reached him. That residual is
  architectural and unchanged: no watcher exists outside GitHub Actions and the
  routine scheduler.
