2026-09-29 — the VRP dispatch stall: 18 days measured (the lead named here was wrong)
Status: open finding, root cause NOT established. This note records the measurement so the next session does not re-derive it. Recorded by the daily audit.
CORRECTION 2026-09-30 (daily audit, prompted by research panel run 31 / PR #194). The mechanism this note originally proposed — "VRP is the only one of the three workflows that declares
workflow_dispatchinputs" — is false, and is retracted below.paper-trading-meanrev.yml:30-35carries an identicaldry_runinputs block and has never been late. Three further corrections are folded in: the late-day count is 9, not 7 (the note's own table always bolded 9); the 09-18 spread close happened at 38 VRP trading days into the 126-day forward gate, not 49 (the pinned evaluatorthales --sleeve vrp forward-gatesays 38, and a merged record retired the 49 figure on 09-18); and PR #174 would have raised 4 of the 9, not all of them. One new observation from 09-30 is appended. The measurement — the table, and the fact that a real timing defect exists — survives all four corrections; only the explanation was wrong.
What is happening
The owner's Mac dispatches the three trading workflows from a single launchd
agent (com.thales.dispatch-trading, Mon–Fri 07:35 PT) through one invocation
of scripts/dispatch_workflows.sh, which loops gh workflow run over
paper-trading.yml, paper-trading-meanrev.yml, paper-trading-vrp.yml in
that order. A healthy day puts all three workflow_dispatch events within about
five seconds of each other.
On 9 of the 18 dispatch days since 2026-09-03, VRP's event arrived minutes to hours after its two siblings. Momentum and meanrev have never been late, not once, on any of those days.
| date | momentum | meanrev | vrp | vrp lag |
|---|---|---|---|---|
| 09-03 | 14:35:06 | 14:35:11 | 14:35:11 | — |
| 09-04 | 14:38:52 | — | 14:46:01 | +7 min |
| 09-07 | 14:40:02 | — | 14:40:06 | — |
| 09-08 | 14:47:25 | — | 14:47:30 | — |
| 09-09 | 14:35:08 | — | 14:35:13 | — |
| 09-10 | 14:35:07 | — | 14:35:13 | — |
| 09-11 | 14:40:21 | — | 16:19:36 | +99 min |
| 09-15 | 14:35:06 | — | 14:35:13 | — |
| 09-16 | 14:37:11 | 14:37:13 | 17:22:59 | +166 min |
| 09-17 | 14:47:56 | 14:47:58 | 19:49:11 | +301 min |
| 09-18 | 14:37:32 | 14:37:35 | 16:36:49 | +119 min |
| 09-21 | 14:35:08 | 14:35:11 | 14:35:13 | — |
| 09-22 | 14:35:05 | 14:35:08 | 14:35:11 | — |
| 09-23 | 14:42:37 | 14:42:39 | 16:38:00 | +115 min |
| 09-24 | 14:37:23 | 14:37:26 | 14:52:56 | +16 min |
| 09-25 | 14:37:23 | 14:37:26 | 17:18:53 | +161 min |
| 09-28 | 14:41:14 | 14:41:17 | 14:41:22 | — |
| 09-29 | 14:41:06 | 14:41:10 | 16:37:00 | +116 min |
| 09-30 | 14:42:13 | 14:42:15 | no dispatch event at all | absent |
Times are created_at (== run_started_at) on workflow_dispatch events, from
the GitHub Actions API. Blank meanrev cells are days not pulled, not missing
runs.
Caveat on four of the nine. 09-23, 09-25 and 09-29 (and 09-30) fall inside the account-level Actions billing refusal that began 09-21: those dispatch events created runs that were refused a runner in 3–11 s and executed nothing. The dispatch lateness is still real and still measured the same way — the event timestamp is written before any runner is involved — but no trading consequence attaches to those days, because nothing traded on any of them. The five pre-outage lates (09-04, 09-11, 09-16, 09-17, 09-18) are the ones that cost something.
It has already cost something. On 2026-09-18 the late VRP run closed a live
SPY put credit spread (data/state/vrp/order_log.jsonl, manage:close) roughly
two hours down the tape from where the schedule intends, on the one live sleeve,
38 VRP trading days into its 126-day forward gate.
2026-09-30: the symptom got worse, not better
On 09-30 momentum and meanrev were dispatched 2 s apart at 14:42:13Z and
14:42:15Z, and VRP was never dispatched at all — no workflow_dispatch run
exists for paper-trading-vrp.yml on 09-30, checked at 22:05Z, 7 h 23 min
after its two siblings. The only VRP run that day is the 19:38:27Z schedule
backup leg. So the failure mode is not merely "arrives late": on at least one
day the third call in the loop produced nothing whatsoever.
That is consistent with the loop's third gh call either hanging past the
launchd job's lifetime or erroring out — and both of those write a
distinguishable line to logs/dispatch.log, which remains the one cheap
measurement nobody has taken.
RETRACTED lead: "VRP is the only one declaring dispatch inputs"
This note originally argued that paper-trading-vrp.yml was the only one of
the three whose workflow_dispatch: carries an inputs: block, so that gh
alone on that call had to resolve an input it was never given — a path that can
block on stdin.
That is factually wrong. paper-trading-meanrev.yml lines 30–35 declare an
identical block:
workflow_dispatch:
inputs:
dry_run:
description: "Preview only — full chain (fetch, signals, selection, safety), submit NOTHING"
type: boolean
default: false
meanrev is dispatched through the same loop, with the same bare
gh workflow run "$wf" --repo "$GH_REPO" call and the same absent -f, and it
has been punctual on 18 of 18 days — including 09-17, the +301 min worst case,
where it was 2 s behind momentum. The inputs block therefore cannot be what
distinguishes VRP. Found by research panel run 31 (PR #194) and re-verified here
by reading both workflow files.
(paper-trading.yml:7 and skew-snapshot.yml:14 do declare a bare
workflow_dispatch: — that part of the original claim was right; it just does
not separate the late sleeve from the punctual one.)
What is left standing
Loop position. VRP is third and last in the dispatcher's argv
(ops/launchd/com.thales.dispatch-trading.plist), and it is the only one of the
three that is ever late. skew-snapshot.yml is fired by a separate launchd
agent, alone and first-in-loop, and has been punctual on all 14 of its dispatch
days in the same window. Every observation in this note is consistent with "the
last call in the loop stalls" and none now requires anything specific to VRP.
That is a correlate, not a mechanism. Nothing measurable from outside the Mac separates "the third call blocks" from "something kills the launchd job after two calls" from "the third call errors and the backup cron covers it".
What settles it
logs/dispatch.logon the Mac for any late day — 09-29 (+116 min) or 09-30 (absent) will do. If thedispatched paper-trading-vrp.ymlline is stamped ~2 h after its two siblings in the same invocation, the call hung. If there is anERROR: dispatch of paper-trading-vrp.yml failedline at 14:42 instead, the call failed. If there is no VRP line at all for 09-30, the job died between the second and third iteration. Three mechanisms, onetail. This is the same ask the 2026-09-17 audit made; it is now 13 days old and is the cheapest item on the board.- Is
HEALTHCHECK_URL_DISPATCHset inconfig/.env? The script beaconsfailon a partial dispatch and, on 09-30, never reaches its beacon call at all if the third iteration hung. No thales-dispatch mail has ever arrived, which is consistent with the beacon being unconfigured — but that is an inference, not a confirmation, and one line from the owner settles it. - The decisive experiment, if the log is inconclusive: reorder the plist's argv so VRP is dispatched first. If VRP still stalls it is something about VRP; if the new last workflow stalls it is loop position. With the inputs hypothesis dead, loop position is the standing prediction.
Proposed remedy (NOT applied — outside the audit's observability rails)
One change to scripts/dispatch_workflows.sh, a no-op on a healthy day:
- if out="$("$GH" workflow run "$wf" --repo "$GH_REPO" 2>&1)"; then
+ if out="$(timeout 60 "$GH" workflow run "$wf" --repo "$GH_REPO" </dev/null 2>&1)"; then
</dev/null removes any blocking-read path; timeout 60 converts a stall into
the loud rc=1 + beacon-fail the script already knows how to report, instead of
a silent multi-hour drift or a silent disappearance. It does not depend on the
retracted hypothesis being right — it makes any stalling call fail loudly on
the minute. The -f dry_run=false half of the original proposal is dropped with
the hypothesis that motivated it. Nothing here touches trading logic, sizing or
any pinned value — but the dispatcher is the mechanism that fires live paper
trades, so it is the owner's change to make, not a routine's.
Detector status
PR #174 (dispatch punctuality alarm) is the right instrument for this class.
Its bar is LATE_AFTER_MINUTES = 90 and it skips runs with a non-decisive
conclusion, so it would have raised 4 of the 9 days here — 09-11, 09-16,
09-17 and 09-18, which are exactly the four pre-outage days over 90 minutes.
It would not have raised 09-04 (+7) or 09-24 (+16), both under its bar, nor the
three outage days whose runs never executed. It also does not, as written, alarm
on 09-30's absent dispatch when a schedule backup leg exists for the same day —
a gap worth checking against the script before merge. The PR has been open and
human-merge-gated since 2026-09-17. Until it merges, this defect is visible only
by hand-diffing created_at timestamps across three workflows — which is how
every one of the ten anomalies here was found.