Thales
← research journal
Sep 29, 2026raw markdown ↗

An internal research document, published verbatim by the automated daily export — not written for an audience, and better for it. All performance discussed is simulated paper trading; nothing here is investment advice.

2026-09-29 — the VRP dispatch stall: 18 days measured (the lead named here was wrong)

Status: open finding, root cause NOT established. This note records the measurement so the next session does not re-derive it. Recorded by the daily audit.

CORRECTION 2026-09-30 (daily audit, prompted by research panel run 31 / PR #194). The mechanism this note originally proposed — "VRP is the only one of the three workflows that declares workflow_dispatch inputs" — is false, and is retracted below. paper-trading-meanrev.yml:30-35 carries an identical dry_run inputs block and has never been late. Three further corrections are folded in: the late-day count is 9, not 7 (the note's own table always bolded 9); the 09-18 spread close happened at 38 VRP trading days into the 126-day forward gate, not 49 (the pinned evaluator thales --sleeve vrp forward-gate says 38, and a merged record retired the 49 figure on 09-18); and PR #174 would have raised 4 of the 9, not all of them. One new observation from 09-30 is appended. The measurement — the table, and the fact that a real timing defect exists — survives all four corrections; only the explanation was wrong.

What is happening

The owner's Mac dispatches the three trading workflows from a single launchd agent (com.thales.dispatch-trading, Mon–Fri 07:35 PT) through one invocation of scripts/dispatch_workflows.sh, which loops gh workflow run over paper-trading.yml, paper-trading-meanrev.yml, paper-trading-vrp.yml in that order. A healthy day puts all three workflow_dispatch events within about five seconds of each other.

On 9 of the 18 dispatch days since 2026-09-03, VRP's event arrived minutes to hours after its two siblings. Momentum and meanrev have never been late, not once, on any of those days.

datemomentummeanrevvrpvrp lag
09-0314:35:0614:35:1114:35:11—
09-0414:38:52—14:46:01+7 min
09-0714:40:02—14:40:06—
09-0814:47:25—14:47:30—
09-0914:35:08—14:35:13—
09-1014:35:07—14:35:13—
09-1114:40:21—16:19:36+99 min
09-1514:35:06—14:35:13—
09-1614:37:1114:37:1317:22:59+166 min
09-1714:47:5614:47:5819:49:11+301 min
09-1814:37:3214:37:3516:36:49+119 min
09-2114:35:0814:35:1114:35:13—
09-2214:35:0514:35:0814:35:11—
09-2314:42:3714:42:3916:38:00+115 min
09-2414:37:2314:37:2614:52:56+16 min
09-2514:37:2314:37:2617:18:53+161 min
09-2814:41:1414:41:1714:41:22—
09-2914:41:0614:41:1016:37:00+116 min
09-3014:42:1314:42:15no dispatch event at allabsent

Times are created_at (== run_started_at) on workflow_dispatch events, from the GitHub Actions API. Blank meanrev cells are days not pulled, not missing runs.

Caveat on four of the nine. 09-23, 09-25 and 09-29 (and 09-30) fall inside the account-level Actions billing refusal that began 09-21: those dispatch events created runs that were refused a runner in 3–11 s and executed nothing. The dispatch lateness is still real and still measured the same way — the event timestamp is written before any runner is involved — but no trading consequence attaches to those days, because nothing traded on any of them. The five pre-outage lates (09-04, 09-11, 09-16, 09-17, 09-18) are the ones that cost something.

It has already cost something. On 2026-09-18 the late VRP run closed a live SPY put credit spread (data/state/vrp/order_log.jsonl, manage:close) roughly two hours down the tape from where the schedule intends, on the one live sleeve, 38 VRP trading days into its 126-day forward gate.

2026-09-30: the symptom got worse, not better

On 09-30 momentum and meanrev were dispatched 2 s apart at 14:42:13Z and 14:42:15Z, and VRP was never dispatched at all — no workflow_dispatch run exists for paper-trading-vrp.yml on 09-30, checked at 22:05Z, 7 h 23 min after its two siblings. The only VRP run that day is the 19:38:27Z schedule backup leg. So the failure mode is not merely "arrives late": on at least one day the third call in the loop produced nothing whatsoever.

That is consistent with the loop's third gh call either hanging past the launchd job's lifetime or erroring out — and both of those write a distinguishable line to logs/dispatch.log, which remains the one cheap measurement nobody has taken.

RETRACTED lead: "VRP is the only one declaring dispatch inputs"

This note originally argued that paper-trading-vrp.yml was the only one of the three whose workflow_dispatch: carries an inputs: block, so that gh alone on that call had to resolve an input it was never given — a path that can block on stdin.

That is factually wrong. paper-trading-meanrev.yml lines 30–35 declare an identical block:

  workflow_dispatch:
    inputs:
      dry_run:
        description: "Preview only — full chain (fetch, signals, selection, safety), submit NOTHING"
        type: boolean
        default: false

meanrev is dispatched through the same loop, with the same bare gh workflow run "$wf" --repo "$GH_REPO" call and the same absent -f, and it has been punctual on 18 of 18 days — including 09-17, the +301 min worst case, where it was 2 s behind momentum. The inputs block therefore cannot be what distinguishes VRP. Found by research panel run 31 (PR #194) and re-verified here by reading both workflow files.

(paper-trading.yml:7 and skew-snapshot.yml:14 do declare a bare workflow_dispatch: — that part of the original claim was right; it just does not separate the late sleeve from the punctual one.)

What is left standing

Loop position. VRP is third and last in the dispatcher's argv (ops/launchd/com.thales.dispatch-trading.plist), and it is the only one of the three that is ever late. skew-snapshot.yml is fired by a separate launchd agent, alone and first-in-loop, and has been punctual on all 14 of its dispatch days in the same window. Every observation in this note is consistent with "the last call in the loop stalls" and none now requires anything specific to VRP.

That is a correlate, not a mechanism. Nothing measurable from outside the Mac separates "the third call blocks" from "something kills the launchd job after two calls" from "the third call errors and the backup cron covers it".

What settles it

  1. logs/dispatch.log on the Mac for any late day — 09-29 (+116 min) or 09-30 (absent) will do. If the dispatched paper-trading-vrp.yml line is stamped ~2 h after its two siblings in the same invocation, the call hung. If there is an ERROR: dispatch of paper-trading-vrp.yml failed line at 14:42 instead, the call failed. If there is no VRP line at all for 09-30, the job died between the second and third iteration. Three mechanisms, one tail. This is the same ask the 2026-09-17 audit made; it is now 13 days old and is the cheapest item on the board.
  2. Is HEALTHCHECK_URL_DISPATCH set in config/.env? The script beacons fail on a partial dispatch and, on 09-30, never reaches its beacon call at all if the third iteration hung. No thales-dispatch mail has ever arrived, which is consistent with the beacon being unconfigured — but that is an inference, not a confirmation, and one line from the owner settles it.
  3. The decisive experiment, if the log is inconclusive: reorder the plist's argv so VRP is dispatched first. If VRP still stalls it is something about VRP; if the new last workflow stalls it is loop position. With the inputs hypothesis dead, loop position is the standing prediction.

Proposed remedy (NOT applied — outside the audit's observability rails)

One change to scripts/dispatch_workflows.sh, a no-op on a healthy day:

-    if out="$("$GH" workflow run "$wf" --repo "$GH_REPO" 2>&1)"; then
+    if out="$(timeout 60 "$GH" workflow run "$wf" --repo "$GH_REPO" </dev/null 2>&1)"; then

</dev/null removes any blocking-read path; timeout 60 converts a stall into the loud rc=1 + beacon-fail the script already knows how to report, instead of a silent multi-hour drift or a silent disappearance. It does not depend on the retracted hypothesis being right — it makes any stalling call fail loudly on the minute. The -f dry_run=false half of the original proposal is dropped with the hypothesis that motivated it. Nothing here touches trading logic, sizing or any pinned value — but the dispatcher is the mechanism that fires live paper trades, so it is the owner's change to make, not a routine's.

Detector status

PR #174 (dispatch punctuality alarm) is the right instrument for this class. Its bar is LATE_AFTER_MINUTES = 90 and it skips runs with a non-decisive conclusion, so it would have raised 4 of the 9 days here — 09-11, 09-16, 09-17 and 09-18, which are exactly the four pre-outage days over 90 minutes. It would not have raised 09-04 (+7) or 09-24 (+16), both under its bar, nor the three outage days whose runs never executed. It also does not, as written, alarm on 09-30's absent dispatch when a schedule backup leg exists for the same day — a gap worth checking against the script before merge. The PR has been open and human-merge-gated since 2026-09-17. Until it merges, this defect is visible only by hand-diffing created_at timestamps across three workflows — which is how every one of the ten anomalies here was found.