December governance package — DRAFT (2026-08-01)
⚠️ DRAFT — NOT REGISTERED — EVERY NUMBER REQUIRES OWNER SIGN-OFF
Panel item C6 (
research/2026-08-01_quant_panel_suggestions.md§3 C6, §6.2): the ladder band, the sunset's retirement-probability acceptance, and the step-down rule are appetite settings the panel cannot make. Nothing here binds anything. This document is deliberately not hash-frozen — freezing a draft would fake a registration. Registration happens only via the §7 sign-off block: each initialed number then lands as a dated, keystone-pinned amendment toconfig/settings.yaml go_live:plus a hash-frozen memo (with theconfig_guard.pyre-pin in the same commit, per house rule). Until then, the only registered December rules remain the 2026-07-05go_live:block and the 2026-07-13 interpretation memo.
Why now: the window is low-power and December's Sharpe outcome is effectively known (report §4.2: live Sharpe −2.34 at n=34 is statistically nothing, and passing needs ~+4.0 annualized over the remaining days) — nothing below can be tuned to flatter. That is precisely what makes this the cheapest moment to write it (report: "closable now precisely because the window is low-power").
The three holes being closed (all at the capital layer, none touched by existing registrations):
go_live.initial_deployment_fraction: 0.25says "review at +63 live td before scaling further" — and names no criteria. A criteria-free review is result-shopping deferred.- Between the retire rule (two consecutive DEGRADED ≈ live Sharpe −2 sustained) and a gate pass (observed ≈ +2.3 at T=126) lies a zombie zone: real money could ride a statistically-dead edge for years inside the −20% kill-switch with no rule ever firing.
- December's report has no economics-per-deployed-dollar line, inviting a post-hoc "the strategy works, deployment is just small" narrative — a reading the red team struck as pre-drafted excuse (§4c below).
1. (a) The scale-up ladder — PROPOSED values
Scope: applies from first real capital (a December PASS) onward. All fractions are of the intended allocation — the dollar figure the owner sets at go-live (itself sign-off item #1; the ladder governs the fraction, not the dollars).
Rungs (proposed): 25% → 50% → 75% → 100%. Entry rung 25% is already
registered (initial_deployment_fraction). One review per rung, at
+63 live trading days on the current rung (the cadence the existing
comment names; ~one quarter).
Rung criteria — ALL must hold over the trailing 63-td rung window:
| # | Criterion | Proposed value | Source / rationale |
|---|---|---|---|
| L1 | Endogenous safety halts (memo §4 taxonomy) | 0 | Carried from max_recent_safety_halts: 0; ambiguity counts against us. ≥3 exogenous halts in the window is itself an operational-reliability failure (memo §4, already registered) → rung fails too. |
| L2 | Real-fill TCA median |slippage| | ≤ 50 bps | The tca_median_slippage_bps_max criterion carried onto REAL fills — the number it was always meant for (paper fills print at NBBO and cannot measure it). |
| L3 | Live-vs-paper tracking error | ≤ 2.0% annualized | TE = std(daily live-minus-paper return) × √252 over the rung window. Derivation of the band: at the 5%/day turnover cap and the 50 bps TCA bound, fill-noise alone implies ~0.8–1.6% annualized; 2.0% tolerates fill noise and catches structural divergence (missed rebalance, wrong book, broken parity). The band is an appetite setting — sign-off item #4. |
| L4 | oos-monitor DEGRADED verdicts dated in-window | 0 | One DEGRADED fails the rung; two consecutive retire outright (§5 precedence). |
Mechanical action (proposed — the red team requires the fraction pinned):
- PASS all four → step UP exactly one rung. No skipping.
- FAIL any → step DOWN exactly one rung (not hold). Rationale: at real-money scale a control failure should shed exposure mechanically; holding preserves the exposure that just failed its own checks — asymmetric to the house fail-closed posture. The alternative (HOLD at the current rung) is listed for the owner's choice — sign-off item #5 selects one.
- Floor: 25%. A failure at the 25% rung HOLDS at 25% — the ladder never retires; retirement belongs exclusively to the DEGRADED rule and the §2 sunset.
- After a step-down, a full 63-td rung window must PASS before re-ascent.
- A rung window with an UNEVALUABLE criterion (dead comparator, §4) is a HOLD — never a step up on missing evidence, and a step down only if an evaluable criterion independently fails.
2. (b) The sunset — PROPOSED values, power stated in print
Rule (proposed): at the first monthly evaluation at or after
378 clean-clock live trading days (~18 months of 21-td months from
oos_monitor.since 2026-06-11 → lands ~2027-12), if ALL of:
- S1. the go-live gate has never passed at any evaluation to date, AND
- S2. the placebo-ensemble percentile is < 60 (the pinned raw percentile — the binding one; C5's beta-hedged companion never gates), AND
- S3. active Sharpe ≤ 0 over the full clean-clock window (active = live daily return − SPY daily return, the memo §2 definition — dodging the beta confound),
then RETIRE: paper lane stops, strategy archived, the never-relitigate covenant applies. This closes the zombie zone: a book that after 18 months has cleared no gate, beats fewer than 60% of random books, and has earned nothing over its own benchmark, no longer gets to consume the shop's scarcest resource (one operator's attention) by merely not being catastrophic.
The power table statement — published and accepted in print (red-team condition: this is signed, not buried):
At 378 td the standard error of an annualized Sharpe is √(252/378) ≈ 0.82. If the strategy's TRUE active Sharpe is +0.3 — a real but modest edge — then P(observed active Sharpe ≤ 0) = Φ(−0.3/0.82) ≈ 36%. The sunset's active-Sharpe leg alone fires on a true-0.3 strategy roughly one time in three. The joint rule fires less often (S1–S3 must all agree), but the owner accepts, in print, that this sunset can retire a true-but-modest edge with material probability. That acceptance is the price of closing the zombie zone; signing it is sign-off item #8. (Conversely: a true-zero strategy escapes the active-Sharpe leg half the time — the sunset is a slow judge in both directions; it is a backstop, not a verdict.)
3. (c) The economics companion — growth per unit of exposure deployed
Metric (report-only, printed at every rung review and every gate evaluation):
GPUE(W) = annualized live total return over window W
÷ mean daily gross exposure over W (gross = Σ|position mv| / equity)
both from the live ledgers; alongside it, always: the mean gross itself, live CAGR, and SPY CAGR over W. If C1's gross-cap amendment is adopted, the denominator tracks whichever go-live diff is currently registered — never a dead counterfactual (A2's red-team condition, mirrored).
SYMMETRIC interpretation — mandated by the red team, both directions carrying both readings:
| Observation | Favorable reading | Equally-available unfavorable/uninformative reading |
|---|---|---|
| GPUE high / rising | the book earns efficiently per deployed dollar | a small denominator amplifying noise; a beta tape lifting a ~0.5-beta book; too few days to mean anything |
| GPUE low / falling | the edge is absent per deployed dollar | a risk-off tape dragging a beta book; deployment constraints binding (Kelly/vol-target), which is the sizing layer working as validated |
The STRUCK reading, named: "the strategy works — deployment is just small" may not be asserted from this number (in either direction). It was struck because this draft is written knowing the live exposure and the modal December outcome — a privileged escape hatch drafted in advance is not analysis. GPUE never upgrades a NOT-PASS, never downgrades a PASS, never feeds any verdict. It exists so the December reader cannot un-know the deployment decomposition.
Disclosure of what was observed at drafting (2026-08-01) — the red-team condition: this draft was written with the frozen panel report in hand and nothing else; no fresh live read was taken for it. Known at drafting, from the report: live gross exposure ≈ 0.419 (= Kelly 0.140 × vol-target scalar 2.996, report §4.1/§4.6); live Sharpe −2.34 at n=34, SE ≈ 2.7 — statistically nothing (§4.2); the negative control +4.5% since 07-13, beating the incumbent (the noise-floor demo working, §4.2); December's modal outcome NOT-PASS (interpretation memo). GPUE itself was NOT computed at drafting. Anyone auditing this draft later judges its symmetry against exactly this disclosure.
4. The comparator assumption — named, with its failure mode
Named assumption: the paper lane keeps trading through the real-money era. It is the comparator for L3 (tracking error) and the context for the companion; the live-vs-paper decomposition is the only instrument that separates "the strategy changed" from "execution changed."
If the paper lane dies (cron failure, broker paper-API sunset, account loss) — proposed handling:
- Gaps ≤ 5 consecutive trading days: tolerated; the dead days are excluded from the TE computation. No backfill ever — paper NBBO fills cannot be honestly reconstructed after the fact.
- Gaps > 5 td in a rung window: L3 is UNEVALUABLE → the rung review result is HOLD (§1 rule: no step up on missing evidence; no step down unless an evaluable criterion independently fails). The next rung window starts when the paper lane resumes.
- Permanent death (e.g. the vendor retires paper trading): L3 cannot be silently dropped — replacing or removing a ladder criterion requires a new owner-signed, dated registration. Until one exists, reviews HOLD. GPUE and the sunset are unaffected (neither needs the paper lane; S2's placebo replays the engine on owned data).
5. Precedence — pre-agreed, and never a reason to loosen DEGRADED
Retirement via two consecutive DEGRADED verdicts supersedes everything here. If it fires first, (a) and (c) are moot — pre-agreed now. The existence of a slower, gentler sunset is never an argument to soften, re-time, or re-interpret the DEGRADED rule (the report's exact caution). Order of authority when rules collide: fleet/sleeve halts → DEGRADED×2 retirement → §2 sunset → ladder. The ladder is the least-privileged rule in the stack; it only ever sizes a book the senior rules have allowed to exist.
6. What this draft deliberately does not do
- No change to any
go_live:value, any keystone, or any live config — drafting is not registering. - No new verdict instrument: L1–L4, S1–S3 are all computed from ledgers
and monitors that already exist (
safetyhalt log + memo taxonomy, TCA log, paper/live equity JSONLs,oos_verdict_history.jsonl,placebo-ensemble,oos-monitorbenchmark block). - No implementation: wiring the rung evaluation into
go-live-gate's report happens after sign-off, as its own change with its own tests.
7. Sign-off block — the owner confirms EXACTLY these numbers
Registration = every line initialed (or amended in the owner's hand), then
committed as a dated keystone-pinned go_live: amendment + hash-frozen
memo. An un-initialed line is not registered; an amended value is the
owner's number, which is the point (report §6.2).
| # | Number / choice | Proposed | Owner initials + date |
|---|---|---|---|
| 1 | Intended allocation (the dollars the fractions apply to) | ______ (owner-only; no proposal) | ______ |
| 2 | Rung fractions | 25% → 50% → 75% → 100% | ______ |
| 3 | Rung review cadence | 63 live td per rung | ______ |
| 4 | L3 tracking-error band (and its definition above) | 2.0% annualized | ______ |
| 5 | Failed-rung action | STEP DOWN one rung (alternative: HOLD) — circle one | ______ |
| 6 | Rung floor / ladder-never-retires | hold at 25%; retirement only via DEGRADED×2 or sunset | ______ |
| 7 | Sunset horizon | 378 clean-clock td (~2027-12) | ______ |
| 8 | Sunset legs + the accepted power number | S1 never-passed ∧ S2 placebo < 60th ∧ S3 active Sharpe ≤ 0; accepting ~36% false-retire on a true-0.3 edge (S3 alone) | ______ |
| 9 | Paper-lane gap tolerance / unevaluable→HOLD | 5 consecutive td | ______ |
| 10 | GPUE rendering points + symmetric-reading mandate (§3 verbatim) | every rung review + every gate evaluation | ______ |
Unchanged and not up for signature here: max_recent_safety_halts: 0,
TCA ≤ 50 bps (L2 restates the registered value onto real fills),
retire_after_consecutive_degraded: 2, the memo's halt taxonomy, and the
2026-12-01 / 126-td / 25% gate itself.
(d) APPENDED 2026-08-02 — exact-money boundary + value-level reconciliation (pre-first-real-dollar)
Source: external design review 2026-08-01 (finding 3, its verified kernel —
archived with corrections at research/2026-08-01_external_design_review.md)
- the standing TECH_DEBT M3 limitation. The review's urgency claim was corrected (reconcile compares order-ID sets and equity DATES, never monetary values, so float drift cannot trip anything today — and reconcile alerts, it does not HALT), but the prospective kernel is right: real dollars deserve exact arithmetic at the money boundary, and M3 says value-level reconciliation should exist anyway.
PROPOSED (owner sign-off with the rest of this draft): before the first real dollar, as ONE designed change with parity tests —
- Execution-layer money arithmetic (order sizing, spend-against-cap, P&L
comparison against broker truth) moves to exact decimal (integer cents or
Decimal). The backtest engine stays float/NumPy by design — vectorized research arithmetic is the correct tool there; the boundary is where broker-truth dollars enter. reconcilegains value-level position/quantity/equity comparison with explicit tolerance bands (closing TECH_DEBT M3), so divergence detection stops depending on ID/date presence alone.
Sequencing note: this lands BEFORE the go-live flip and AFTER December's verdict is read (a live-path arithmetic change during the evaluation window would churn the thing being measured).