The reconcile false-drift at the window edge (2026-09-15, corrected 2026-09-16)
Class: detector defect (observability tier). Found by: daily audit, 2026-09-15.
Corrected and re-fixed by: daily audit 2026-09-16, after panel run 24 (PR #172)
reviewed the first fix pre-merge. Fixed by: PR from branch
audit/reconcile-window-edge-2026-09-15.
Read the 2026-09-16 correction at the bottom before reusing anything here. The original diagnosis got the mechanism right and the period of the defect wrong, and the first fix it proposed (
days + 2) would have left the largest observed instance still flagged. Both are corrected below; the intermediate claims are kept, struck through in prose, because the way the model failed is the reusable part.
What was observed
Paper Trading (meanrev) went red on 09-11 (both legs), 09-14 and 09-15 (both legs)
— six red runs, six 🔴 "sleeve trade job failure" emails. Every failure was the same
step, running after trading had already completed and committed:
thales portfolio reconcile --days 14
251 broker orders missing locally:
2026-09-01 buy MUR qty=0.386012868 id=1031b2a1… status=filled
... and 231 more
🔴 BROKER/LOCAL BOOK DRIFT — investigate
What it turned out to be
The orders are not missing. Eight of eight flagged broker order ids were looked up in
the committed data/state/meanrev/order_log.jsonl and all eight are present — MUR,
STX, C, VTR, MTSI, NCLH, NI, MCHP — each recorded locally as 2026-08-31 while the
broker stamps it 2026-09-01.
reconcile() compares two sets built with different boundary semantics:
- the broker side fetches from an exact instant,
after = midnight(as_of) - days, and dates each order by the broker's ownsubmitted_at; - the local side calls
state.read_orders(last_n_days=days), which keeps rows whose date string is>= as_of - days.
The local run-date can lag the broker's submitted_at by a calendar day (an order
submitted near or after the close lands in the broker's next session). At the window's
lower edge that skew drops the local rows out of the read while their broker twins stay
in range — so the whole boundary batch reads as "missing locally".
The defect is an asymmetry, not a new bug. The opposite direction
(missing_in_broker_orders) was given a 2-day window-edge tolerance on 2026-06-11 for
this exact root cause (boundary = (after + timedelta(days=2)).date(), reconcile.py).
The missing_in_local_orders direction never got the mirror.
Why it did not clear on its own (as diagnosed 09-15 — partly wrong, see the correction)
Two previous audits predicted it would clear the next day. Both were wrong. The 09-15
diagnosis explained that as: the boundary moves with the window, so on any day the
sleeve traded, the batch dated exactly as_of - days reappears as a phantom orphan —
concluding it could never self-heal. That conclusion is false, and 2026-09-16
falsified it: see correction §2. What actually has to sit on the edge is a displaced
batch, not any batch.
The mechanism does make an exact, falsifiable prediction — the flagged date is always
as_of - 14 — and the record confirms it on all three observed runs:
| run date | predicted flagged date | observed | phantom orphans |
|---|---|---|---|
| 2026-09-11 | 2026-08-28 | 2026-08-28 | 55 |
| 2026-09-14 | 2026-08-31 | 2026-08-31 | 222 |
| 2026-09-15 | 2026-09-01 | 2026-09-01 | 251 |
What it cost
- Alert fatigue on the loudest channel. Six false 🔴 emails in three sessions, in the same envelope a genuine trading failure arrives in.
- A silenced liveness beacon. The healthchecks ping in
paper-trading-meanrev.ymlis gatedif: ${{ success() }}, so a red reconcile withholds it.thales-meanrevwas DOWN from 09-11 because of a false alarm — meaning a genuinely missed meanrev run would have looked identical.
That second cost is the serious one: a detector defect had disabled a detector.
(Qualified 2026-09-16 by panel run 24's BCN-1: the _alert.yml job runs on any
non-success trade job and did email on all three red days, and digest_exceptions
forces a same-day email on no_run_sleeves. So a genuinely missed run would NOT have
looked identical — identical only on the healthchecks channel. The beacon was still
wrongly held down for four days.)
The fix (as shipped, 2026-09-16)
The membership set for the broker→local direction is read unwindowed. The question that direction asks is "was this broker order ever written locally?" — a set-membership question with no time window in it. The window only ever belonged to the broker fetch, which bounds how far back we ask the broker to look, and that side is deliberately left alone: narrowing it, or skipping broker orders near its edge (the mirror of the local→broker tolerance), is the direction that would hide real drift, because it removes broker orders from the check instead of adding ids to the match set. (Corrected in review 2026-09-26: this sentence first said "widening", which is backwards. Asking the broker to look further back can only add orders to check. The fetch window is now pinned by a test.)
Unwindowing the membership set cannot mask a silent loss, for two independent reasons, each pinned by its own negative control:
- an order that was never written locally is absent at every window width;
- matching is by exact
order_id, so no widening lets an unrelated local row stand in for a missing one (control: a local record sharing the broker order's symbol, side and date but differing in id must still leave the broker order flagged).
A third control pins the other direction — the local→broker tolerance stays bounded at 2 days, because that is the side where unbounded really would hide drift.
Cost is nil: cli.py already reads the whole order log unwindowed two calls earlier
(1.1 MB on the busiest sleeve).
as_of is also threaded into the windowed local read, which had been falling back to
date.today(). No production caller passes as_of, so this is behaviour-neutral
today — which is exactly what made it a latent trap.
Tests: 6 functions / 8 cases: the skew-width pin (parametrized ×3), 3 negative controls, the end-to-end edge batch (a clean batch plus one real loss, pinning the fix and the catch together), and a pin on the broker fetch window (added in review 2026-09-26). Verification, run three ways:
local_ids read | midweek 1d | weekend 3d | long weekend 4d | controls |
|---|---|---|---|---|
last_n_days=days (pre-fix) | ❌ fails | ❌ fails | ❌ fails | all pass |
last_n_days=days + 2 (first proposal) | ✅ passes | ❌ fails | ❌ fails | all pass |
| unwindowed (shipped) | ✅ | ✅ | ✅ | all pass |
tests/test_execution/ — 579 passed.
Correction, 2026-09-16 — two things the 09-15 diagnosis got wrong
1. The skew is "next trading SESSION", not "next calendar day" — so days + 2 was too narrow
An order submitted at or after the close lands in the broker's next session: one calendar day mid-week, three across a weekend (Friday run → Monday stamp), four before a Monday holiday. Replaying the window arithmetic against the live record:
| run | broker window start | local twin's date | days + 2 cutoff | repaired by +2? |
|---|---|---|---|---|
| 2026-09-11 | 08-28 | 08-27 (Thu) | 08-26 | ✅ |
| 2026-09-14 | 08-31 | 08-28 (Fri) | 08-29 | ❌ 222 still flagged |
| 2026-09-15 | 09-01 | 08-31 (Mon) | 08-30 | ✅ |
So the fix as first written would have repaired two of the three cases in its own
evidence table and left the largest one — the 222-record batch — still firing,
while landing a record saying the defect was closed. The table in "The fix" above is
the empirical confirmation: under days + 2 the weekend and long-weekend cases fail.
This was already on the record before the fix was written: panel run 23 posted exactly this critique on PR #166 at 2026-09-15 13:35Z, including the prescription to add headroom or make the read trading-day-aware and to build the regression test from the 08-28 → 08-31 shape. PR #171 was opened nine hours later and took neither half. Panel run 24 (PR #172) caught it pre-merge.
Any fixed width is a guess at the market calendar. Unwindowing is not a wider guess; it removes the question.
2. "It can never self-heal" was wrong — and today proved it
The 09-15 note argued that because the boundary moves with the window and meanrev trades daily, "there is always a batch sitting on the edge". That does not follow. What has to sit on the edge is not any batch but a displaced one — a batch whose broker stamp rolled into the next session, which only happens when the run itself fires at or after the 20:00Z close. Those existed only during the late-August displaced-cron period (meanrev ran 23:51Z on 08-28 and 20:32Z on 08-31); every run from 09-01 onward fired at 18:07Z or earlier.
So the alarm cleared on its own the moment the 14-day window rolled past 08-31's batch.
2026-09-16: both meanrev legs green, reconcile clean, with PR #171 still unmerged —
and thales-meanrev came back UP at 14:40Z after 4 d 16 h down. Panel run 24
pre-registered this outcome before the 14:40Z run ("0 or 4 phantom orphans, most likely
green") and was right.
That green is self-healing luck, not evidence the fix is unneeded. The defect is
dormant, not absent: it returns whenever a run is displaced past the close and then
resurfaces exactly days later. The displaced-cron regime that produced it recurred as
recently as three weeks ago.
Standing lessons
- The 2026-06-11 fix repaired one direction of a symmetric bug and left its mirror live for three months. When a boundary-tolerance fix lands on one side of a two-way comparison, the other side is a defect until a test says otherwise.
- A tolerance expressed in calendar days, guarding a market-calendar phenomenon, is wrong at every width. Ask whether the comparison needs a window at all before picking one.
- A control that passes at the width you shipped and at the width you should have
shipped is not a control. The first version's
test_tolerance_is_bounded_at_two_dayspassed at+2, passed at+3, and failed only at+4— so it did not pin what its name claimed, and its only discriminating power was to forbid the correct fix. The fixture helper was weekday-blind by construction, so it could not express a weekend at all. Third time in three weeks that a fixture agreed with the half of a bug that was fixed. - A prediction of self-healing that fails twice means the model is wrong — and so does one that succeeds unexpectedly. Both directions are evidence. The 09-15 note correctly rejected "it will clear tomorrow" and then over-corrected into "it can never clear", which the very next session falsified.