# 2026-09-29 — the VRP dispatch stall: 18 days measured (the lead named here was wrong)

**Status:** open finding, root cause NOT established. This note records the
measurement so the next session does not re-derive it. Recorded by the daily
audit.

> **CORRECTION 2026-09-30 (daily audit, prompted by research panel run 31 /
> PR #194).** The mechanism this note originally proposed — "VRP is the only
> one of the three workflows that declares `workflow_dispatch` inputs" — is
> **false**, and is retracted below. `paper-trading-meanrev.yml:30-35` carries
> an identical `dry_run` inputs block and has never been late. Three further
> corrections are folded in: the late-day count is **9**, not 7 (the note's own
> table always bolded 9); the 09-18 spread close happened at **38** VRP trading
> days into the 126-day forward gate, not 49 (the pinned evaluator
> `thales --sleeve vrp forward-gate` says 38, and a merged record retired the
> 49 figure on 09-18); and PR #174 would have raised **4** of the 9, not all of
> them. One new observation from 09-30 is appended. The measurement — the
> table, and the fact that a real timing defect exists — survives all four
> corrections; only the explanation was wrong.

## What is happening

The owner's Mac dispatches the three trading workflows from a single launchd
agent (`com.thales.dispatch-trading`, Mon–Fri 07:35 PT) through one invocation
of `scripts/dispatch_workflows.sh`, which loops `gh workflow run` over
`paper-trading.yml`, `paper-trading-meanrev.yml`, `paper-trading-vrp.yml` in
that order. A healthy day puts all three `workflow_dispatch` events within about
five seconds of each other.

On **9** of the 18 dispatch days since 2026-09-03, VRP's event arrived minutes
to hours after its two siblings. Momentum and meanrev have never been late, not
once, on any of those days.

| date | momentum | meanrev | vrp | vrp lag |
|---|---|---|---|---|
| 09-03 | 14:35:06 | 14:35:11 | 14:35:11 | — |
| 09-04 | 14:38:52 | — | 14:46:01 | **+7 min** |
| 09-07 | 14:40:02 | — | 14:40:06 | — |
| 09-08 | 14:47:25 | — | 14:47:30 | — |
| 09-09 | 14:35:08 | — | 14:35:13 | — |
| 09-10 | 14:35:07 | — | 14:35:13 | — |
| 09-11 | 14:40:21 | — | 16:19:36 | **+99 min** |
| 09-15 | 14:35:06 | — | 14:35:13 | — |
| 09-16 | 14:37:11 | 14:37:13 | 17:22:59 | **+166 min** |
| 09-17 | 14:47:56 | 14:47:58 | 19:49:11 | **+301 min** |
| 09-18 | 14:37:32 | 14:37:35 | 16:36:49 | **+119 min** |
| 09-21 | 14:35:08 | 14:35:11 | 14:35:13 | — |
| 09-22 | 14:35:05 | 14:35:08 | 14:35:11 | — |
| 09-23 | 14:42:37 | 14:42:39 | 16:38:00 | **+115 min** |
| 09-24 | 14:37:23 | 14:37:26 | 14:52:56 | **+16 min** |
| 09-25 | 14:37:23 | 14:37:26 | 17:18:53 | **+161 min** |
| 09-28 | 14:41:14 | 14:41:17 | 14:41:22 | — |
| 09-29 | 14:41:06 | 14:41:10 | 16:37:00 | **+116 min** |
| **09-30** | 14:42:13 | 14:42:15 | **no dispatch event at all** | **absent** |

Times are `created_at` (== `run_started_at`) on `workflow_dispatch` events, from
the GitHub Actions API. Blank meanrev cells are days not pulled, not missing
runs.

**Caveat on four of the nine.** 09-23, 09-25 and 09-29 (and 09-30) fall inside
the account-level Actions billing refusal that began 09-21: those dispatch
events created runs that were refused a runner in 3–11 s and executed nothing.
The *dispatch* lateness is still real and still measured the same way — the
event timestamp is written before any runner is involved — but no trading
consequence attaches to those days, because nothing traded on any of them. The
five pre-outage lates (09-04, 09-11, 09-16, 09-17, 09-18) are the ones that cost
something.

**It has already cost something.** On 2026-09-18 the late VRP run closed a live
SPY put credit spread (`data/state/vrp/order_log.jsonl`, `manage:close`) roughly
two hours down the tape from where the schedule intends, on the one live sleeve,
**38** VRP trading days into its 126-day forward gate.

## 2026-09-30: the symptom got worse, not better

On 09-30 momentum and meanrev were dispatched 2 s apart at 14:42:13Z and
14:42:15Z, and **VRP was never dispatched at all** — no `workflow_dispatch` run
exists for `paper-trading-vrp.yml` on 09-30, checked at 22:05Z, 7 h 23 min
after its two siblings. The only VRP run that day is the 19:38:27Z schedule
backup leg. So the failure mode is not merely "arrives late": on at least one
day the third call in the loop produced nothing whatsoever.

That is consistent with the loop's third `gh` call either hanging past the
launchd job's lifetime or erroring out — and both of those write a
distinguishable line to `logs/dispatch.log`, which remains the one cheap
measurement nobody has taken.

## RETRACTED lead: "VRP is the only one declaring dispatch inputs"

This note originally argued that `paper-trading-vrp.yml` was the only one of
the three whose `workflow_dispatch:` carries an `inputs:` block, so that `gh`
alone on that call had to resolve an input it was never given — a path that can
block on stdin.

**That is factually wrong.** `paper-trading-meanrev.yml` lines 30–35 declare an
identical block:

```yaml
  workflow_dispatch:
    inputs:
      dry_run:
        description: "Preview only — full chain (fetch, signals, selection, safety), submit NOTHING"
        type: boolean
        default: false
```

meanrev is dispatched through the same loop, with the same bare
`gh workflow run "$wf" --repo "$GH_REPO"` call and the same absent `-f`, and it
has been punctual on 18 of 18 days — including 09-17, the +301 min worst case,
where it was 2 s behind momentum. The inputs block therefore cannot be what
distinguishes VRP. Found by research panel run 31 (PR #194) and re-verified here
by reading both workflow files.

(`paper-trading.yml:7` and `skew-snapshot.yml:14` do declare a bare
`workflow_dispatch:` — that part of the original claim was right; it just does
not separate the late sleeve from the punctual one.)

## What is left standing

**Loop position.** VRP is third and last in the dispatcher's argv
(`ops/launchd/com.thales.dispatch-trading.plist`), and it is the only one of the
three that is ever late. `skew-snapshot.yml` is fired by a *separate* launchd
agent, alone and first-in-loop, and has been punctual on all 14 of its dispatch
days in the same window. Every observation in this note is consistent with "the
last call in the loop stalls" and none now requires anything specific to VRP.

That is a correlate, not a mechanism. Nothing measurable from outside the Mac
separates "the third call blocks" from "something kills the launchd job after
two calls" from "the third call errors and the backup cron covers it".

## What settles it

1. **`logs/dispatch.log` on the Mac** for any late day — 09-29 (+116 min) or
   09-30 (absent) will do. If the `dispatched paper-trading-vrp.yml` line is
   stamped ~2 h after its two siblings *in the same invocation*, the call hung.
   If there is an `ERROR: dispatch of paper-trading-vrp.yml failed` line at
   14:42 instead, the call failed. If there is **no VRP line at all** for 09-30,
   the job died between the second and third iteration. Three mechanisms, one
   `tail`. This is the same ask the 2026-09-17 audit made; it is now 13 days
   old and is the cheapest item on the board.
2. **Is `HEALTHCHECK_URL_DISPATCH` set in `config/.env`?** The script beacons
   `fail` on a partial dispatch and, on 09-30, never reaches its beacon call at
   all if the third iteration hung. No thales-dispatch mail has ever arrived,
   which is consistent with the beacon being unconfigured — but that is an
   inference, not a confirmation, and one line from the owner settles it.
3. **The decisive experiment, if the log is inconclusive:** reorder the plist's
   argv so VRP is dispatched first. If VRP still stalls it is something about
   VRP; if the new last workflow stalls it is loop position. With the inputs
   hypothesis dead, loop position is the standing prediction.

## Proposed remedy (NOT applied — outside the audit's observability rails)

One change to `scripts/dispatch_workflows.sh`, a no-op on a healthy day:

```diff
-    if out="$("$GH" workflow run "$wf" --repo "$GH_REPO" 2>&1)"; then
+    if out="$(timeout 60 "$GH" workflow run "$wf" --repo "$GH_REPO" </dev/null 2>&1)"; then
```

`</dev/null` removes any blocking-read path; `timeout 60` converts a stall into
the loud `rc=1` + beacon-fail the script already knows how to report, instead of
a silent multi-hour drift or a silent disappearance. It does not depend on the
retracted hypothesis being right — it makes *any* stalling call fail loudly on
the minute. The `-f dry_run=false` half of the original proposal is dropped with
the hypothesis that motivated it. Nothing here touches trading logic, sizing or
any pinned value — but the dispatcher is the mechanism that fires live paper
trades, so it is the owner's change to make, not a routine's.

## Detector status

`PR #174` (dispatch punctuality alarm) is the right instrument for this class.
Its bar is `LATE_AFTER_MINUTES = 90` and it skips runs with a non-decisive
conclusion, so it would have raised **4** of the 9 days here — 09-11, 09-16,
09-17 and 09-18, which are exactly the four pre-outage days over 90 minutes.
It would not have raised 09-04 (+7) or 09-24 (+16), both under its bar, nor the
three outage days whose runs never executed. It also does not, as written, alarm
on 09-30's *absent* dispatch when a schedule backup leg exists for the same day —
a gap worth checking against the script before merge. The PR has been open and
human-merge-gated since 2026-09-17. Until it merges, this defect is visible only
by hand-diffing `created_at` timestamps across three workflows — which is how
every one of the ten anomalies here was found.
