Two of the VRP forward gate's reported numbers were wrong (daily audit 2026-09-18)
Both found by reconciling one sleeve's numbers against each other on a single evening. Neither changed a decision's direction; both changed what the record said, and one of them ran on the owner's daily channel for weeks.
Recording the pair together because they share a shape worth naming: a reporting layer maintaining its own copy of a definition that a pinned evaluator already owns. Neither was a bug in a calculation. Both were a second definition, drifting.
1. The daily email's forward-gate clock was 11 trading days ahead of the gate
thales --sleeve vrp forward-gate (the pinned evaluator) said 38/126. The
fleet email said 49/126. One line of fleet_digest.py — "gate_days": len(equity_hist) — was responsible, via two compounding errors:
- it counted from the first equity row (2026-07-13, activation) rather than the
pinned clean clock
oos_monitor.since= 2026-07-27, which was reset once, deliberately and on the record, by the fidelity amendment (research/2026-07-25_vrp_decision_fidelity_amendment.md) — 10 of the 11 days; - it counted equity MARKS where the gate counts RETURNS — the 11th.
Why it is worse than an off-by-N. The whole purpose of resetting since on
the record was that the earlier returns must not count toward this gate. The
reporting layer counted them anyway. So a pre-registered clock reset was being
silently undone on the one surface the owner reads every morning — the exact
failure that pinning a clock in advance exists to prevent.
It propagated, too: the 2026-09-17 daily audit reported "day 48/126" because it quoted the digest rather than the evaluator. The correct figure that day was 37.
Fixed in PR #177: the digest calls the same resolve_since +
load_live_returns the gate uses, so there is one definition. Pinned by a test
on a fixture that spans the reset, plus a control on the fixture itself so the
pin cannot pass degenerately.
2. The pessimistic twin's headline was ~2x the honest figure, and its cohort count was wrong
The gate's companion line read -$100.00 (mid basis $14.00; 6 opens / 6 closes).
twin_net_credit summed every row in the ledger, which overstated twice over:
- it included the four rows 2026-07-13..2026-07-31 that the ledger's own
annotationrecord marks basis-contaminated (panel A1: the two bases priced from different quote fetches; 07-31 printed a negative drag, impossible from a single quote). The frozen v2 spec §5 excludes exactly those "by construction"; - it read a null as a zero. The 07-13 open logged
pessimistic_credit: null(occ-key bug);float(x or 0)scored that as zero credit received while still charging its paired close $52 in full. That asymmetry alone is $52 of the apparent $-100.
Honest figure on the spec's basis: -$52.00 pessimistic against +$1.00 mid, over 4 complete cohorts.
The conclusion is unchanged and in fact starker — the twin is negative while the mid curve is positive, which is precisely v2 §6's "a pass that exists only on the mid curve is not a pass". Nothing scored moved; the twin is a companion note, never a criterion.
What did move is the clock. v2's deferred drag constant D needs ≥ 8
qualifying cohorts and the spec says "the count is the clock". The line said 6;
the qualifying count is 4. A reader tracking v2 activation would have thought
it two cohorts away when it is four.
Fixed in PR #176, with the raw all-rows sum still reported so the correction reconciles against what earlier reports printed.
The generalisable lesson
Both defects are the same species as the 2026-09-14 dark-routine finding and the
2026-09-15 reconcile window edge: a consumer re-deriving a producer's
definition instead of asking it. The tell is a local expression that looks
obviously right (len(equity_hist); sum every row) standing in for a rule
defined elsewhere with conditions attached — a reset clock, an annotation, a
schema version.
Cheap standing check for future audits, which is how both of these surfaced: when two artefacts describe the same object, read them side by side. The VRP section of the digest and the VRP gate output are two descriptions of one sleeve. Diffing them took minutes and found both.
A corollary for this routine specifically: quote the evaluator, not the email. The 09-17 audit's wrong figure came from trusting the reporting layer it exists to audit.
Amendment — 2026-09-21, research panel run 27 (append only; nothing above rewritten)
The sentence "That asymmetry alone is $52 of the apparent $-100" attributes the correction to the wrong mechanism. Measured against the real ledger by decomposing the two fixes:
MAIN (old): mid 14.00 pess -100.00 opens 6 closes 6
NULL-FIX ONLY: mid -39.00 pess -100.00 opens 5 closes 6 unpriced 1
ANNOTATION ONLY: mid 1.00 pess -52.00 opens 4 closes 4
BOTH (as shipped): mid 1.00 pess -52.00 n_cohorts 4
n_excluded_annotated 4 n_excluded_unpriced 0
The annotation exclusion does 100% of the -$100 → -$52 move. The null-fix
alone leaves pessimistic at exactly -$100.00 — dropping a row that
contributed 0 under float(x or 0) changes nothing, because its paired close
is still charged $52 in full either way. And in the shipped code the null path
fires zero times (n_excluded_unpriced = 0), because the 2026-07-13 open is
itself inside the annotation's annotates date list and the annotation branch
continues first.
The claim is literally true of the raw arithmetic but is presented as one of two reasons the headline "overstated twice over", which reads as additive. It is not: the two fixes are nested, not independent.
Nothing else in this memo moves. The clock decomposition was reproduced
exactly and independently (marks_since_activation 49 / marks_since_clean 39 / returns_since_clean 38 as of 09-18 → gap 11 = reset 10 + marks-vs-returns 1),
as were -$52.00 / +$1.00 / 4 qualifying cohorts, and the 09-17 figures
(48 printed, 37 correct).
Why the correction is worth making rather than letting stand. It records a fact the original reading hides: the null-drop code has never been exercised on live data. Every row it would have caught was already excluded for a different reason. That matters for whoever reviews that branch next — its behaviour on a genuinely unannotated null is untested by the live replay, and measured forward it is wrong (it drops the row from the mid total too, and still charges the paired close alone). Raised pre-merge on #176.
This is the same class as the memo's own closing lesson: a quantity that looks obviously right standing in for a rule defined elsewhere. Here it was an attribution that looked obviously right standing in for a decomposition nobody had run.