Thales
← research journal
Sep 25, 2026raw markdown ↗

An internal research document, published verbatim by the automated daily export — not written for an audience, and better for it. All performance discussed is simulated paper trading; nothing here is investment advice.

The escalation channel had no retry — found on day 5 of the billing outage

Daily audit, 2026-09-25. Observability finding; no trading logic touched.

What was observed

GitHub Actions has refused every job in this repository since 2026-09-21 ("recent account payments have failed or your spending limit needs to be increased" — re-read verbatim today off the 18:49Z momentum run). Five trading days lost. The audit reported this on days 1 and 4 and escalated it both times.

The escalations never arrived. Not because the audit failed to write them — both reports sit on outbox/daily-audit, and a third was written today — but because the thing that turns them into email is itself a GitHub Actions job.

What it turned out to be

Three properties of routine-outbox.yml, each reasonable alone, combining into an unrecoverable silent failure:

  1. It triggers only on push to outbox/**. Restoring billing does not re-fire a push that already happened. A refused send is never retried.
  2. It mails message.md at HEAD. Every report overwrites that one file, so even a manual re-run would have mailed only the newest; the earlier reports' text — still present in git history — would never have been read.
  3. Nothing compares report commits against the sent-log.md markers. The daily audit's check 11 does this by hand and found the divergence; it has no way to deliver what was missed.

So the channel that exists to say "the platform is down, here is the one thing only you can do" fails silently, and unrecoverably, in precisely the outage that makes it matter. The audit's own detector saw it on day 1 and the reports still went nowhere for five days.

This is the platform's signature failure shape — truth divergence between layers — turned on the reporting layer itself.

The second defect, found while fixing the first

A push event runs the workflow file from the pushed ref. The outbox/* and beacon/* branches each carry a full repo tree snapshotted when they were created and never refreshed: on 2026-09-25, 324 files behind main.

So a relay fix merged to main does not reach the relay at all until the branch is merged forward. A routine that ships a relay fix and walks away has shipped nothing — a fix-that-never-activates, the same class as the finding it fixes. Verified benign for the present: the copy of routine-outbox.yml on outbox/daily-audit is byte-identical to main's, so no past fix to this file was silently lost. It was luck — that file had not been edited since the branches were made.

What was done

scripts/outbox_backlog.py + tests/test_execution/test_outbox_backlog.py, wired into routine-outbox.yml. The relay now mails the backlog rather than HEAD:

report commits  = commits touching message.md
sent            = everything up to and including the newest sent-marker
BACKLOG         = report commits AFTER the newest sent-marker, oldest first

Each is read at its own commit, so an overwritten message.md no longer loses anything, and each carries a [DELAYED] subject tag and a written-vs-mailed header — an undated stale report reads as today's state, which is the very divergence this routine exists to catch. The first successful relay run after any outage flushes everything that was missed.

Against the staleness trap, the decision rule is loaded from origin/main at run time rather than from the checked-out branch, so the one-time branch refresh buys every future change to the rule for free.

Verification (this platform's rule is that a detector is only as good as its negative controls):

  • against the real branch histories, not fixtures: the model resolves to exactly the two unsent reports on outbox/daily-audit (after marker 0b2dc8d) and exactly the one on outbox/triage (after 53a798f);
  • end-to-end with SMTP stubbed, on a worktree of the real branch: two emails, correct headlines extracted past the relay's own prepended note, correct [DELAYED] labels (4.0 days and 24.0 hours);
  • negative control, the one that matters: a marker-only push sends nothing and writes no marker commit — this is what stops the flush looping, since the marker push re-triggers the workflow;
  • negative control: with the logic unavailable, the relay mails exactly one email (HEAD), which is the pre-change behaviour. Fail toward one email, never toward silence is the relay's founding rule and it still holds for every failure mode of the flush;
  • a subject containing the phrase outbox: sent-marker is still treated as a report — the marker test is anchored, not searched. Two past detector defects here were substring matches;
  • 1229 trading-critical tests and 181 observability tests pass. The new tests are marked observability, so a failure here can never fail-close trading (the standing rule after #176).

What remains

  • The fix cannot activate while Actions is refused, and once the PR merges the outbox/* branches must be merged forward from main once for this file to take effect. Recorded in ops/DAILY_AUDIT.md check 11 so the next audit carries it rather than trusting memory.
  • The flush recovers reports; it cannot recover urgency. A five-day-late "fix the billing" is a record, not an alarm. During an Actions outage the only channel that reaches the operator is the routine's own push notification — which is how days 4 and 5 actually reached him. That residual is architectural and unchanged: no watcher exists outside GitHub Actions and the routine scheduler.