DISMISSED — CUT-1 — the cutover's gate reports an unreachable box as a postponed gate, so an uncurable failure is retried seven more times; considered as a loudness row and REJECTED on scope and self-retirement
status: dismissed at adversarial review, panel run 36 (2026-10-07) · class considered: loudness & safety (G2) · every mechanism claim is TRUE and was verified from the repo plus the owner's own inbox; the row is killed because the defect lives in one script that removes itself on 2026-10-16, and the operational fact it carries was delivered the same morning by notification rather than by a queue slot
What was considered
The tbx-win cutover agent fired on schedule on its FIRST_DAY, Tue 2026-10-06
at 18:00 PT, and postponed. The owner's email (sent 2026-10-07T01:44:19Z,
subject ⏸ thales tbx-win cutover postponed: a gate did not hold (retrying tomorrow 18:00)) carries gate 1's value verbatim:
=== tbx-win cutover --scheduled ===
gate 1 (practice): Connection timed out during banner exchange | Connection to UNKNOWN port 65535 timed out
gate 2 (docs PR): #206 mergeable=MERGEABLE
NOT SWITCHING: a gate does not hold. Nothing was changed; the Mac + GitHub keep running live.
Gate 1 did not return a verdict. It returned an SSH transport error. The proposal was that the gate cannot distinguish three states with three different remedies, and that the retry policy added by #214 is correct for one of them and actively misleading for another.
The mechanism claims — all verified, recorded so nobody re-derives them
-
scripts/win/cutover.sh:41—ps1() { ssh -o BatchMode=yes -o ConnectTimeout=30 "$HOST" "$@" 2>&1 | ... }. The2>&1folds SSH's stderr into the captured value, so a transport failure becomes the content of$gatesrather than a non-zero status. -
scripts/win/cutover.sh:79,92-93— gate 1 isgrep -q "^trading CLEAN"andgrep -q "^skew CLEAN"against that value. Any text that is not apractice_gate.pyverdict fails the grep. This is fail-closed and correct; it is also lossy — "box not ready" and "box not reachable" produce the identical branch. -
scripts/win/cutover.sh:96→retry_or_give_up(:63-69) → the email subject⏸ … postponed: a gate did not hold (retrying tomorrow 18:00). Waiting cures an unclean practice day (#214's rolling window is exactly that fix). Waiting does not cure an sshd that will not complete a banner exchange. -
scripts/win/practice_gate.py:21— the gate's inputs areC:\thales\home\logs\{trading,skew}.log, on the box. Confirmed: no cloud session can evaluate gate 1, so PR #221's claim that gate 1 "reads the box's own practice logs, not the Mac's dispatches" is TRUE. -
scripts/win/session.py:13,297,332-333— in practice modepersistislive_onlyandpush_checkis a--dry-runagainst a ref the comment says is "never created". A practising box leaves zero trace in the repo, by design. This is what the 2026-10-05 triage email meant by "after tomorrow's cutover, nothing watches the box". -
scripts/win/cutover.sh:33-34—FIRST_DAY="2026-10-06",LAST_DAY="2026-10-16", whose own comment is "stop retrying here: the Actions allowance is forecast to run out ~10-20". Eight weekday evenings remain (10-07, 08, 09, 12, 13, 14, 15, 16);[[ "$(date +%F)" < "$LAST_DAY" ]]makes Fri 2026-10-16 the give-up evening, so seven more "retrying tomorrow" emails then one "gave up". -
The 09-28 audit's forecast, on the record in PR #191: the Actions allowance exhausts again 2026-10-19/20 (window 10-19 → 10-23), and "no NORTHSTAR §6 row carries it". Mon 2026-10-19 is three days after the cutover's give-up date.
-
Realized harm, and it is the reason this was considered at all: the 2026-10-06 daily audit's single escalation was "Tonight's cutover is the fix. Tell me if it postpones", and the remedy it pre-wrote for the postponement branch was
launchctl list | grep thales— on the Mac. That is a correct repair for the skew-dispatch regression and it is the wrong machine for the cause that actually fired. The audit's decision tree has no branch for "the box is unreachable", because the email it told the owner to read does not say so in its subject. -
Secondary leg, measured: gate 1 has no overall deadline.
ConnectTimeout=30bounds the TCP connect only, and nothing bounds the SSH banner exchange, which is where this failure sat. The agent is documented as firing at 18:00 PT (ops/windows-host/README.md:248-250); the email left at 18:44 PT, so the run hung roughly 44 minutes on one unanswered SSH. Harmless here — the script changes nothing before its gates — but it is the same shape asbuilt/skw-1-skew-client-no-deadline.md, which this shop judged worth building. Recorded so the reopen conditions below can cite a measured number rather than a worry.
Why it was killed anyway
- The generalization does not hold, and it was the row's load-bearing
claim. The row was drafted as "a pattern, not an instance": a remote
command's stderr folded into a value a decision is then grepped out of.
Tested rather than asserted — the idiom appears in exactly two places
in the repo,
cutover.sh:41andseed-from-mac.sh:30(a one-time seeding script that gates nothing), and the one place that looks similar in the trading path,dispatch_workflows.sh:75, branches on the exit code (if out="$(… 2>&1)"; then) and is therefore correct. Nothing in the box's steady-state operation uses the channel at all: after cutover the box runs its own Task Scheduler, andboot.pySSHes to GitHub, not from the Mac. So the defect is one script, not a class. - That script deletes itself on 2026-10-16. A row whose subject self-retires in nine days, against a queue at its stated 12-item cap on a Monday that already forces an overflow dismissal (HIRE-1's automatic reopen), is a slot spent on something that expires before the triage → implementer → human-merge path can plausibly deliver it. The same arithmetic killed a verified finding in run 35.
^scripts/is in the auto-merge guard's DENY_PATTERNS (scripts/automerge_guard.py:63), so the fix needs a human merge whichever routine authors it. A queue row therefore buys no mechanical advantage over the fix stated verbatim in the run's report — which is where it is stated.- The operational fact was already delivered, eight hours before the channel that would otherwise carry it. This run pushed the finding to the owner's phone at ~13:5xZ on 2026-10-07; the daily audit fires 22:00Z and had already pre-committed to checking this exact branch ("Tomorrow I check whether the capture landed, at what hour, and whether the cutover switched"). The gap the panel closed was timing, not knowledge, and timing is not a research row.
What is NOT claimed, so no later sweep treats it as established
- Whether the box's deploy key is registered is NOT-VERIFIABLE from here.
The 2026-10-02 practice email (
🔴 [PRACTICE] trading (tbx-win) failure) says "the push check failed: the box could not push to GitHub", andops/windows-host/README.md:215makes registering the key an owner step. Whether that was since done cannot be checked from a routine container —PUSH_CHECK_REFis never created, so the attempt leaves no trace. Stated as unknown, deliberately: DSP-1's dismissal was corrected in #221 for exactly this error class. - Whether the Mac is asleep at midday is still open. The cutover agent firing at 18:00 PT on 10-06 proves the Mac's launchd was alive that evening; it says nothing about noon, which is when the skew dispatch runs. It does narrow "the launchd subsystem is dead" to "this agent, or this time of day".
- Cross-project evidence, labelled as such and not built on. The owner's
inbox carries
tbx/remote weekly check — all ok(2026-10-03T06:20:15Z,reports[at]tbx.report, written with the@broken because the public-export scrub gate BLOCKS any bare email address — seeresearch/2026-10-07_a_records_pr_froze_the_public_site.md) reportinghost:reachable OK tbx-win answered. That is the other project's watcher on the same physical box, it is weekly, and its next run falls after four more futile cutover evenings. Coupling thales's gate to another project's monitor is a cross-project design decision and is the owner's to make, not a routine's.
The fix, stated here so the row is not needed to carry it
Four lines in scripts/win/cutover.sh, replacing the single lossy branch at
:92-96 with a transport check before it:
if ! grep -qE "^(trading|skew) (CLEAN|NOT-READY)" <<<"$gates"; then
say "NOT SWITCHING: gate 1 did not run — the box did not answer: ${gates//$'\n'/ | }"
[ "$MODE" = "--check" ] && exit 1
retry_or_give_up "🔴 thales tbx-win cutover BLOCKED: the box is unreachable (not a practice-day failure — waiting will not fix this)"
fi
The point is the subject line, not the retry: an owner who reads "⏸ postponed (retrying tomorrow)" correctly does nothing, and an owner who reads "🔴 BLOCKED: the box is unreachable" goes and looks at the box.
Reopens if
- (a) the cutover's life is extended —
LAST_DAYmoves past 2026-10-16, or a successor script inherits the same gate — since the defect's horizon then exceeds the cost of fixing it; or - (b) a second consumer gates a decision on the text of a remote command with stderr folded in, which converts claim 1 from an instance into the pattern this row could not show it was; or
- (c) a second cutover evening postpones on a transport error and the
owner, acting on the "⏸ postponed (retrying tomorrow)" subject or on the
10-06 audit's
launchctl listremedy, spends another cycle on the Mac while the box is the blocker — i.e. the mis-subjecting demonstrably costs a second day; or - (d) the cutover gives up on 2026-10-16 with the box never having been reached, AND the 2026-10-19/20 Actions outage arrives with the Mac's skew dispatch still broken — the conjunction this row forecast, at which point the question is no longer this script's email wording but the fleet's dispatch architecture, and it belongs in NORTHSTAR §6 with an evaluator.