# The reconcile false-drift at the window edge (2026-09-15, corrected 2026-09-16)

**Class:** detector defect (observability tier). **Found by:** daily audit, 2026-09-15.
**Corrected and re-fixed by:** daily audit 2026-09-16, after panel run 24 (PR #172)
reviewed the first fix pre-merge. **Fixed by:** PR from branch
`audit/reconcile-window-edge-2026-09-15`.

> **Read the 2026-09-16 correction at the bottom before reusing anything here.** The
> original diagnosis got the mechanism right and the *period* of the defect wrong, and
> the first fix it proposed (`days + 2`) would have left the largest observed instance
> still flagged. Both are corrected below; the intermediate claims are kept, struck
> through in prose, because the way the model failed is the reusable part.

## What was observed

`Paper Trading (meanrev)` went red on 09-11 (both legs), 09-14 and 09-15 (both legs)
— six red runs, six 🔴 "sleeve trade job failure" emails. Every failure was the same
step, running *after* trading had already completed and committed:

```
thales portfolio reconcile --days 14
  251 broker orders missing locally:
    2026-09-01  buy  MUR   qty=0.386012868  id=1031b2a1…  status=filled
    ... and 231 more
🔴 BROKER/LOCAL BOOK DRIFT — investigate
```

## What it turned out to be

The orders are not missing. Eight of eight flagged broker order ids were looked up in
the committed `data/state/meanrev/order_log.jsonl` and **all eight are present** — MUR,
STX, C, VTR, MTSI, NCLH, NI, MCHP — each recorded locally as **2026-08-31** while the
broker stamps it **2026-09-01**.

`reconcile()` compares two sets built with different boundary semantics:

- the broker side fetches from an exact **instant**, `after = midnight(as_of) - days`,
  and dates each order by the broker's own `submitted_at`;
- the local side calls `state.read_orders(last_n_days=days)`, which keeps rows whose
  **date string** is `>= as_of - days`.

The local run-date can lag the broker's `submitted_at` by a calendar day (an order
submitted near or after the close lands in the broker's next session). At the window's
lower edge that skew drops the local rows out of the read while their broker twins stay
in range — so the whole boundary batch reads as "missing locally".

**The defect is an asymmetry, not a new bug.** The *opposite* direction
(`missing_in_broker_orders`) was given a 2-day window-edge tolerance on 2026-06-11 for
this exact root cause (`boundary = (after + timedelta(days=2)).date()`, reconcile.py).
The `missing_in_local_orders` direction never got the mirror.

## Why it did not clear on its own (as diagnosed 09-15 — *partly wrong, see the correction*)

Two previous audits predicted it would clear the next day. Both were wrong. The 09-15
diagnosis explained that as: *the boundary moves with the window, so on any day the
sleeve traded, the batch dated exactly `as_of - days` reappears as a phantom orphan* —
concluding it could never self-heal. **That conclusion is false**, and 2026-09-16
falsified it: see correction §2. What actually has to sit on the edge is a *displaced*
batch, not any batch.

The mechanism does make an exact, falsifiable prediction — the flagged date is always
`as_of - 14` — and the record confirms it on all three observed runs:

| run date | predicted flagged date | observed | phantom orphans |
|---|---|---|---|
| 2026-09-11 | 2026-08-28 | 2026-08-28 | 55 |
| 2026-09-14 | 2026-08-31 | 2026-08-31 | 222 |
| 2026-09-15 | 2026-09-01 | 2026-09-01 | 251 |

## What it cost

1. **Alert fatigue on the loudest channel.** Six false 🔴 emails in three sessions, in
   the same envelope a genuine trading failure arrives in.
2. **A silenced liveness beacon.** The healthchecks ping in
   `paper-trading-meanrev.yml` is gated `if: ${{ success() }}`, so a red reconcile
   withholds it. `thales-meanrev` was DOWN from 09-11 **because of a false alarm** —
   meaning a genuinely missed meanrev run would have looked identical.

That second cost is the serious one: a detector defect had disabled a detector.

*(Qualified 2026-09-16 by panel run 24's BCN-1: the `_alert.yml` job runs on any
non-success trade job and did email on all three red days, and `digest_exceptions`
forces a same-day email on `no_run_sleeves`. So a genuinely missed run would NOT have
looked identical — identical only on the healthchecks channel. The beacon was still
wrongly held down for four days.)*

## The fix (as shipped, 2026-09-16)

The **membership set** for the broker→local direction is read **unwindowed**. The
question that direction asks is *"was this broker order ever written locally?"* — a
set-membership question with no time window in it. The window only ever belonged to the
broker fetch, which bounds how far back we ask the broker to look, and that side is
deliberately left alone: *narrowing* it, or skipping broker orders near its edge (the
mirror of the local→broker tolerance), is the direction that would hide real drift,
because it removes broker orders from the check instead of adding ids to the match set.
*(Corrected in review 2026-09-26: this sentence first said "widening", which is
backwards. Asking the broker to look further back can only add orders to check. The
fetch window is now pinned by a test.)*

Unwindowing the membership set cannot mask a silent loss, for two independent reasons,
each pinned by its own negative control:

1. an order that was never written locally is absent at **every** window width;
2. matching is by exact `order_id`, so no widening lets an unrelated local row stand in
   for a missing one (control: a local record sharing the broker order's symbol, side
   and date but differing in id must still leave the broker order flagged).

A third control pins the **other** direction — the local→broker tolerance stays bounded
at 2 days, because that is the side where unbounded really would hide drift.

Cost is nil: `cli.py` already reads the whole order log unwindowed two calls earlier
(1.1 MB on the busiest sleeve).

`as_of` is also threaded into the windowed local read, which had been falling back to
`date.today()`. No production caller passes `as_of`, so this is behaviour-neutral
today — which is exactly what made it a latent trap.

Tests: 6 functions / 8 cases: the skew-width pin (parametrized ×3), 3 negative
controls, the end-to-end edge batch (a clean batch plus one real loss, pinning the fix
and the catch together), and a pin on the broker fetch window (added in review
2026-09-26).
**Verification, run three ways:**

| `local_ids` read | midweek 1d | weekend 3d | long weekend 4d | controls |
|---|---|---|---|---|
| `last_n_days=days` (pre-fix) | ❌ fails | ❌ fails | ❌ fails | all pass |
| `last_n_days=days + 2` (first proposal) | ✅ passes | ❌ **fails** | ❌ **fails** | all pass |
| unwindowed (shipped) | ✅ | ✅ | ✅ | all pass |

`tests/test_execution/` — **579 passed**.

---

## Correction, 2026-09-16 — two things the 09-15 diagnosis got wrong

### 1. The skew is "next trading SESSION", not "next calendar day" — so `days + 2` was too narrow

An order submitted at or after the close lands in the broker's next *session*: one
calendar day mid-week, **three** across a weekend (Friday run → Monday stamp), four
before a Monday holiday. Replaying the window arithmetic against the live record:

| run | broker window start | local twin's date | `days + 2` cutoff | repaired by `+2`? |
|---|---|---|---|---|
| 2026-09-11 | 08-28 | 08-27 (Thu) | 08-26 | ✅ |
| **2026-09-14** | **08-31** | **08-28 (Fri)** | **08-29** | ❌ **222 still flagged** |
| 2026-09-15 | 09-01 | 08-31 (Mon) | 08-30 | ✅ |

So the fix as first written would have repaired two of the three cases in its own
evidence table and left the **largest** one — the 222-record batch — still firing,
while landing a record saying the defect was closed. The table in "The fix" above is
the empirical confirmation: under `days + 2` the weekend and long-weekend cases fail.

This was **already on the record before the fix was written**: panel run 23 posted
exactly this critique on PR #166 at 2026-09-15 13:35Z, including the prescription to
add headroom *or* make the read trading-day-aware and to build the regression test from
the 08-28 → 08-31 shape. PR #171 was opened nine hours later and took neither half.
Panel run 24 (PR #172) caught it pre-merge.

Any **fixed** width is a guess at the market calendar. Unwindowing is not a wider
guess; it removes the question.

### 2. "It can never self-heal" was wrong — and today proved it

The 09-15 note argued that because the boundary moves with the window and meanrev
trades daily, *"there is always a batch sitting on the edge"*. That does not follow.
What has to sit on the edge is not any batch but a **displaced** one — a batch whose
broker stamp rolled into the next session, which only happens when the run itself fires
at or after the 20:00Z close. Those existed only during the late-August displaced-cron
period (meanrev ran 23:51Z on 08-28 and 20:32Z on 08-31); every run from 09-01 onward
fired at 18:07Z or earlier.

So the alarm cleared on its own the moment the 14-day window rolled past 08-31's batch.
**2026-09-16: both meanrev legs green, reconcile clean, with PR #171 still unmerged** —
and `thales-meanrev` came back UP at 14:40Z after 4 d 16 h down. Panel run 24
pre-registered this outcome before the 14:40Z run ("0 or 4 phantom orphans, most likely
green") and was right.

That green is **self-healing luck, not evidence the fix is unneeded**. The defect is
dormant, not absent: it returns whenever a run is displaced past the close and then
resurfaces exactly `days` later. The displaced-cron regime that produced it recurred as
recently as three weeks ago.

### Standing lessons

1. The 2026-06-11 fix repaired one direction of a symmetric bug and left its mirror
   live for three months. **When a boundary-tolerance fix lands on one side of a
   two-way comparison, the other side is a defect until a test says otherwise.**
2. **A tolerance expressed in calendar days, guarding a market-calendar phenomenon, is
   wrong at every width.** Ask whether the comparison needs a window at all before
   picking one.
3. **A control that passes at the width you shipped and at the width you should have
   shipped is not a control.** The first version's `test_tolerance_is_bounded_at_two_days`
   passed at `+2`, passed at `+3`, and failed only at `+4` — so it did not pin what its
   name claimed, and its only discriminating power was to *forbid* the correct fix. The
   fixture helper was weekday-blind by construction, so it could not express a weekend
   at all. Third time in three weeks that a fixture agreed with the half of a bug that
   was fixed.
4. **A prediction of self-healing that fails twice means the model is wrong — and so
   does one that succeeds unexpectedly.** Both directions are evidence. The 09-15 note
   correctly rejected "it will clear tomorrow" and then over-corrected into "it can
   never clear", which the very next session falsified.
