# Two of the VRP forward gate's reported numbers were wrong (daily audit 2026-09-18)

Both found by reconciling one sleeve's numbers against each other on a single
evening. Neither changed a decision's *direction*; both changed what the record
said, and one of them ran on the owner's daily channel for weeks.

Recording the pair together because they share a shape worth naming: **a
reporting layer maintaining its own copy of a definition that a pinned evaluator
already owns.** Neither was a bug in a calculation. Both were a second
definition, drifting.

---

## 1. The daily email's forward-gate clock was 11 trading days ahead of the gate

`thales --sleeve vrp forward-gate` (the pinned evaluator) said **38/126**. The
fleet email said **49/126**. One line of `fleet_digest.py` — `"gate_days":
len(equity_hist)` — was responsible, via two compounding errors:

- it counted from the first equity row (2026-07-13, activation) rather than the
  pinned clean clock `oos_monitor.since` = **2026-07-27**, which was **reset
  once, deliberately and on the record**, by the fidelity amendment
  (`research/2026-07-25_vrp_decision_fidelity_amendment.md`) — 10 of the 11 days;
- it counted equity MARKS where the gate counts RETURNS — the 11th.

**Why it is worse than an off-by-N.** The whole purpose of resetting `since` on
the record was that the earlier returns must not count toward this gate. The
reporting layer counted them anyway. So a pre-registered clock reset was being
**silently undone on the one surface the owner reads every morning** — the exact
failure that pinning a clock in advance exists to prevent.

It propagated, too: the 2026-09-17 daily audit reported "day 48/126" because it
quoted the digest rather than the evaluator. The correct figure that day was 37.

Fixed in **PR #177**: the digest calls the same `resolve_since` +
`load_live_returns` the gate uses, so there is one definition. Pinned by a test
on a fixture that *spans the reset*, plus a control on the fixture itself so the
pin cannot pass degenerately.

## 2. The pessimistic twin's headline was ~2x the honest figure, and its cohort count was wrong

The gate's companion line read `-$100.00 (mid basis $14.00; 6 opens / 6 closes)`.
`twin_net_credit` summed every row in the ledger, which overstated twice over:

- it **included the four rows 2026-07-13..2026-07-31** that the ledger's own
  `annotation` record marks basis-contaminated (panel A1: the two bases priced
  from *different* quote fetches; 07-31 printed a negative drag, impossible from
  a single quote). The frozen v2 spec §5 excludes exactly those **"by
  construction"**;
- it **read a null as a zero**. The 07-13 open logged `pessimistic_credit: null`
  (occ-key bug); `float(x or 0)` scored that as *zero credit received* while
  still charging its paired close $52 in full. That asymmetry alone is $52 of the
  apparent $-100.

Honest figure on the spec's basis: **-$52.00 pessimistic against +$1.00 mid, over
4 complete cohorts.**

**The conclusion is unchanged and in fact starker** — the twin is negative while
the mid curve is positive, which is precisely v2 §6's "a pass that exists only on
the mid curve is not a pass". Nothing scored moved; the twin is a companion note,
never a criterion.

**What did move is the clock.** v2's deferred drag constant `D` needs **≥ 8
qualifying cohorts** and the spec says "the count is the clock". The line said 6;
the qualifying count is **4**. A reader tracking v2 activation would have thought
it two cohorts away when it is four.

Fixed in **PR #176**, with the raw all-rows sum still reported so the correction
reconciles against what earlier reports printed.

---

## The generalisable lesson

Both defects are the same species as the 2026-09-14 dark-routine finding and the
2026-09-15 reconcile window edge: **a consumer re-deriving a producer's
definition instead of asking it.** The tell is a local expression that *looks*
obviously right (`len(equity_hist)`; sum every row) standing in for a rule
defined elsewhere with conditions attached — a reset clock, an annotation, a
schema version.

Cheap standing check for future audits, which is how both of these surfaced:
**when two artefacts describe the same object, read them side by side.** The VRP
section of the digest and the VRP gate output are two descriptions of one sleeve.
Diffing them took minutes and found both.

A corollary for this routine specifically: **quote the evaluator, not the email.**
The 09-17 audit's wrong figure came from trusting the reporting layer it exists to
audit.

---

## Amendment — 2026-09-21, research panel run 27 (append only; nothing above rewritten)

**The sentence "That asymmetry alone is $52 of the apparent $-100" attributes
the correction to the wrong mechanism.** Measured against the real ledger by
decomposing the two fixes:

```
MAIN (old):         mid  14.00   pess -100.00   opens 6  closes 6
NULL-FIX ONLY:      mid -39.00   pess -100.00   opens 5  closes 6  unpriced 1
ANNOTATION ONLY:    mid   1.00   pess  -52.00   opens 4  closes 4
BOTH (as shipped):  mid   1.00   pess  -52.00   n_cohorts 4
                    n_excluded_annotated 4      n_excluded_unpriced 0
```

The annotation exclusion does **100%** of the `-$100 → -$52` move. The null-fix
**alone leaves pessimistic at exactly `-$100.00`** — dropping a row that
contributed `0` under `float(x or 0)` changes nothing, because its paired close
is still charged $52 in full either way. And in the shipped code the null path
**fires zero times** (`n_excluded_unpriced = 0`), because the 2026-07-13 open is
itself inside the annotation's `annotates` date list and the annotation branch
`continue`s first.

The claim is literally true of the *raw* arithmetic but is presented as one of
two reasons the headline "overstated twice over", which reads as additive. It is
not: the two fixes are nested, not independent.

**Nothing else in this memo moves.** The clock decomposition was reproduced
exactly and independently (`marks_since_activation 49 / marks_since_clean 39 /
returns_since_clean 38` as of 09-18 → gap 11 = reset 10 + marks-vs-returns 1),
as were `-$52.00` / `+$1.00` / 4 qualifying cohorts, and the 09-17 figures
(48 printed, 37 correct).

**Why the correction is worth making rather than letting stand.** It records a
fact the original reading hides: **the null-drop code has never been exercised on
live data.** Every row it would have caught was already excluded for a different
reason. That matters for whoever reviews that branch next — its behaviour on a
genuinely unannotated null is untested by the live replay, and measured
forward it is wrong (it drops the row from the *mid* total too, and still charges
the paired close alone). Raised pre-merge on #176.

This is the same class as the memo's own closing lesson: a quantity that looks
obviously right standing in for a rule defined elsewhere. Here it was an
attribution that looked obviously right standing in for a decomposition nobody
had run.
