# 2026-09-21 — GitHub refused every job in the repo over billing; nothing traded, and nothing said so

**Status:** INCIDENT, live at the time of writing. Root cause identified and
quoted verbatim from GitHub; the fix is the owner's (a billing page), not a
code change. Recorded by the daily audit.

## What happened

No sleeve traded on Monday 2026-09-21. `main` carries no commit for the day —
no state, no site export, no capture. Every GitHub Actions run created for this
repository since some point over the weekend has failed, including workflows
that share no code with each other.

The last successful run of any kind was **2026-09-18 22:37:09Z**. The first
failure was **2026-09-21 13:43:47Z**. Nothing ran on 09-19 or 09-20, so the
state flipped somewhere inside that weekend window.

Today's inventory, from the Actions API: **zero successes.** The three trading
sleeves were dispatched on time from the owner's Mac at 14:35:08 / :11 / :13Z
and all three were refused; their backup crons at 19:40 / 19:42 / 19:44Z were
refused; the fleet digest at 20:07Z was refused; the routine beacon relay and
the outbox mailer were refused; auto-merge was refused. A push made at 22:18Z
while writing this was refused too, so the block was still in force then.

## The cause, in GitHub's own words

Every failed job carries the same annotation, which the audit read at
`GET /repos/MaksimDan/thales/check-runs/{id}/annotations`:

> "The job was not started because recent account payments have failed or your
> spending limit needs to be increased. Please check the 'Billing & plans'
> section in your settings"

This repository is **private**, so its Actions minutes are billed against the
account's allowance rather than being free. Either the allowance is spent with
a spending limit that forbids overage, or a payment on file failed. Both land
on the same page, and neither is visible from inside the repo.

## The signature, so the class is cheap to recognise next time

This is the third distinct way this platform's jobs have died, and it looks
like neither of the first two:

| | runner | duration | steps | logs | conclusion |
|---|---|---|---|---|---|
| ordinary failure | assigned | full | present, one red | present | `failure` |
| runner loss (2026-08-06) | `runner_name: ""` | ~15 min | none | none | `cancelled` |
| **platform refusal (today)** | `runner_id: 0`, `runner_name: ""` | **3–11 s** | none | **404** | `failure` |

`get_workflow_run_usage` closes it: `billable.UBUNTU.total_ms: 0`. Nothing ran,
so nothing was billed — and a 404 on logs is not a tooling problem, it is the
absence of a log stream that was never opened.

The discriminator worth remembering is that **the annotations endpoint names
the cause outright.** One API call, no inference. Three separate routines today
(panel, weekly audit, daily audit) saw the failures; the queue-triage routine
was the one that went and read the annotation, and it is the only place the
actual sentence appears.

## The part that matters more than the outage: the alarm inverted

The operator's inbox today contained exactly three Thales messages, all
healthchecks.io **DOWN** notices — `thales-panel` (17:00Z),
`thales-weekly-audit` (20:00Z), `thales-triage` (22:00Z).

**All three of those routines ran perfectly.** Each pushed its beacon commit on
schedule — `beacon/panel` 13:47:13Z, `beacon/weekly-audit` 14:22:15Z,
`beacon/triage` 16:17:10Z — and each opened its PR. What failed was the
*relay*: `routine-beacon.yml` is a GitHub Actions workflow, so it was refused
along with everything else, and the ping never reached healthchecks.io.

Meanwhile the thing that actually broke — three sleeves not trading — produced
**no alert at all**. The `alert / email` job inside each trading run is itself
an Actions job in the same refused run (jobs 106377069686 / 106377075520 /
106377093282). The fleet digest that would have reported "SLEEVE DID NOT RUN"
was refused. The heartbeat monitor, the dead-man's switch built for exactly
this, is an Actions job too.

So the alarm layer said, loudly and three times, that the research routines
were dead — which was false — and said nothing whatsoever about the trading
platform, which was true. A reader who trusted the inbox would have spent the
evening on the wrong problem.

This is the structural residual §4b of `ops/DAILY_AUDIT.md` already named — *no
in-band alarm can be present during an outage that takes out the whole band* —
except that it now has a sharper edge. The escape hatch §4b nominated was
"healthchecks.io and these cloud routines". Healthchecks fired, and it fired
**wrong**, because the routine beacons reach it through the very band that was
down. `routine-beacon.yml`'s own header records the two earlier lessons that
forced the relay design (routine containers cannot reach `hc-ping.com`; they
cannot use `repository_dispatch` either). Today adds the third: the relay
inherits Actions' availability, so a routine DOWN is only evidence about the
routine when Actions is up.

**The Actions-independent arbiter exists and costs one command.** Routines push
their beacon commits with a contents write, which needs no runner:

```bash
git log -1 --date=iso --pretty='%ad | %s' origin/beacon/<routine>
```

A beacon branch that moved on schedule under a DOWN email means the relay
failed, not the routine. That check is now written into the daily audit's
procedure (§4c, check 5).

## What this cost

- **A trading day on all three sleeves.** Momentum's daily volatility check and
  its drawdown kill-switch did not evaluate — the committed kill-switch state
  is still stamped `last_date: 2026-09-18`. Had the book breached its −20%
  threshold today, nothing would have fired.
- **A decision on the one live sleeve.** VRP closed its spread on 09-18 and is
  flat; the pinned rule re-enters on the next run. Whatever it would have done
  today, it did not do, inside a 126-trading-day forward clock currently at 38.
- **A forward-capture day that cannot be refetched.** The skew snapshot never
  ran. Option skew, chains, IV term structure, borrow flags and the PiT
  snapshot are point-in-time streams; 09-21 is simply absent from them. The
  FINRA short-volume leg is backfillable; the rest is not.
- **Both report channels.** The outbox relay is an Actions workflow, so the
  triage routine's report and this audit's report were written and not mailed.
  Both routines fell back to a phone notification.

## What was NOT affected

The cloud routines themselves, which run on Anthropic's infrastructure and
reach GitHub through the API: panel, weekly audit and triage all completed and
opened PRs (#179, #180, #181). The repository is not disabled, the scheduler is
alive (it created today's runs), and no workflow file has changed since
2026-09-07. This is an account-level refusal, not a defect in the tree.

---

## Days 2–4 (2026-09-22 → 09-24): the outage held, and the watcher went dark

Appended by the daily audit of 2026-09-24. The paragraphs above were written on
day 1 and are left as they stood; this section is what four days changed.

**The refusal is unchanged and still account-level.** Read again today off the
`trade` job of the 18:36Z momentum run (check-run 107777545034), the annotation
is the same sentence, word for word:

> "The job was not started because recent account payments have failed or your
> spending limit needs to be increased. Please check the 'Billing & plans'
> section in your settings"

Every run on 09-22, 09-23 and 09-24 died in 3–11 seconds with zero steps —
schedule-triggered and Mac-dispatched alike. `main` has carried no commit since
2026-09-18. The Mac dispatcher is healthy and kept firing on time all four days
(`workflow_dispatch` at 14:35Z, 14:37Z, 14:42Z, 14:52Z); it had nothing to
dispatch onto.

**Cost, revised from one day to four.** Four trading days on all three sleeves.
Four un-refetchable forward-capture days (`thales captures` reports all seven
time-gated streams at 4 days stale and not accruing). Four days in which
momentum's drawdown kill-switch did not evaluate. And four days off both
forward clocks — which has a dated consequence: momentum stands at 68 of the
126 live trading days its go-live gate requires, so even if trading resumed on
2026-09-25 the 126th day now falls on **2026-12-16**, past the gate's
`not_before: 2026-12-01`. Each further lost day moves that date one further.

### The finding: a dark cloud routine has no detector at all

Day 1's lesson was that the alarm layer pointed *away* from the failure. Days
2–4 sharpened it into something worse.

The daily audit did not run on 09-22 or on 09-23. Neither did the research
panel on any of the three days, nor the implementer on 09-23. The evidence is
Actions-independent and therefore trustworthy: the beacon branches are written
by `git push` from the routine container, which needs no runner, and not one of
them moved after 2026-09-21 22:23Z. Two weekday audits are simply missing.

Nothing anywhere reported this, and the reason is structural:

- The fleet digest's dark-routine exception is an Actions job — refused.
- The healthchecks beacons relay through `routine-beacon.yml` — also an Actions
  job, so every check was *already* DOWN for relay reasons. `thales-daily-audit`
  went DOWN at 2026-09-22 02:00Z and stayed there; healthchecks notifies on the
  transition, so the genuinely-dark days produced no new mail and could not have.
- The cloud routines are each other's only out-of-band witness, and none of them
  watches the others.

So during an Actions outage the beacon channel loses precisely the distinction
it exists to make: **"the routine ran and the relay died" and "the routine never
ran" present identically.** On 09-21 it was the first. On 09-22 and 09-23 it was
the second. Nothing in the system could tell them apart.

**And §4c, as written on day 1, made that worse.** It taught the next reader
that a routine DOWN under a refused-Actions state is a relay artefact — true
that day, and the owner was told so explicitly. Two days later the same signal
was real. An audit that trains its operator to dismiss a class of alarm owes
that operator the symmetric case in the same breath; day 1's text gave the
discriminator (`git log origin/beacon/<routine>`) but framed only one of its two
outcomes as a finding.

The correction is check 12 below — cheap, self-referential, and the one thing
this routine can verify about itself without trusting any layer that an outage
can take down.

### Residual — narrowed 2026-09-25

No detector inside this repository can catch a dark cloud routine during an
Actions outage, because every in-repo detector needs a runner. Check 12 catches
it *retrospectively*, on the next audit that does run — which is a real
improvement over nothing, and is not the same as catching it live.

The original wording of this residual — that closing the live gap needs a
watcher outside **both** Actions and the routine scheduler — was **too strong**,
and panel run 28 (2026-09-25) demonstrated it by doing the thing it declared
impossible: it queried the routine control plane, which is Actions-independent,
and established that the implementer *did* fire on 09-23 and died after 16
seconds. Beacon evidence cannot distinguish that from never firing. So the
fire-vs-die question this audit escalated to the owner on 09-24 is answerable
in-band, at least for that routine, and the escalation should not have been
framed as requiring him.

What remains genuinely open is narrower and worth stating exactly, because the
two routines disagree about their own reach: the daily audit's session could
**not** reproduce that query (2026-09-25 — `CronList` returns only jobs created
by `CronCreate` within the calling session, not the account's routine roster).
So the capability exists but is not uniformly reachable, and which tool the
panel used is not recorded in its report. **Next action, cheap and not the
owner's:** have whichever routine can reach the control plane write down the
exact call, so check 12 can escalate from a retrospective branch-date read to a
live liveness query. Only if no routine can reach it does this become the
owner-level architecture decision originally claimed.

## Amendment 2026-09-26: which of the two causes it was

The owner read the account's billing page on 2026-09-22 (Settings → Billing &
plans, personal account). It resolves the either/or in "The cause, in GitHub's
own words" to its **first** branch: the allowance was spent. No payment failed.

- Plan: **GitHub Free**, no payment method on file, no payment history, so
  there was no payment that could fail.
- September (1–30) gross metered usage **$12.06**, of which included usage
  covered **$12.06**: the month's whole free allowance, consumed. By repository:
  thales $5.37, tesla-toolbox-companion $5.07, docket $1.24, three others
  $0.38. The allowance is shared by every private repo on the account, so thales
  did not exhaust it alone.
- With no payment method, overage cannot be billed, so once the included
  minutes were gone GitHub refused every job. The annotation's wording ("recent
  account payments have failed **or** your spending limit needs to be
  increased") covers both cases; this was the second.

Consequences worth keeping:

1. **It ends by itself on 2026-10-01**, when the next period's allowance
   starts, hours before that day's 14:35Z dispatch (momentum's monthly
   selection). No owner click is needed for that.
2. **The owner has decided to stay on the free tier** (2026-09-22). So the
   remedy is not "restore billing"; it is "fit inside the free allowance across
   all repos". At September's pace thales used roughly 65–70 billable minutes
   per trading day, about 1,450 in a 22-session month. The largest single waste
   is the backup `schedule` leg of each trading workflow, which re-runs the whole
   job (about 14 minutes for momentum) after the Mac dispatch leg already did
   the day's work. Every job is also rounded up to a whole minute, so the many
   few-second relay and guard jobs cost a minute each.
3. **It will recur whenever the allowance runs out**, most likely late in a
   month. The signature above identifies it from the repo side; the billing
   page is the only place the numbers are. Routines should stop escalating
   "check Billing & plans" as the fix; the useful escalation is how close the
   month is to the allowance.
