# December governance package — DRAFT (2026-08-01)

> ## ⚠️ DRAFT — NOT REGISTERED — EVERY NUMBER REQUIRES OWNER SIGN-OFF
>
> Panel item C6 (`research/2026-08-01_quant_panel_suggestions.md` §3 C6,
> §6.2): the ladder band, the sunset's retirement-probability acceptance,
> and the step-down rule are **appetite settings the panel cannot make**.
> Nothing here binds anything. This document is deliberately **not**
> hash-frozen — freezing a draft would fake a registration. Registration
> happens only via the §7 sign-off block: each initialed number then lands
> as a dated, keystone-pinned amendment to `config/settings.yaml go_live:`
> plus a hash-frozen memo (with the `config_guard.py` re-pin in the same
> commit, per house rule). Until then, the only registered December rules
> remain the 2026-07-05 `go_live:` block and the 2026-07-13 interpretation
> memo.

**Why now:** the window is low-power and December's Sharpe outcome is
effectively known (report §4.2: live Sharpe −2.34 at n=34 is statistically
nothing, and passing needs ~+4.0 annualized over the remaining days) —
**nothing below can be tuned to flatter**. That is precisely what makes
this the cheapest moment to write it (report: "closable now precisely
because the window is low-power").

**The three holes being closed** (all at the capital layer, none touched by
existing registrations):

1. `go_live.initial_deployment_fraction: 0.25` says "review at +63 live td
   before scaling further" — and names **no criteria**. A criteria-free
   review is result-shopping deferred.
2. Between the retire rule (two consecutive DEGRADED ≈ live Sharpe −2
   sustained) and a gate pass (observed ≈ +2.3 at T=126) lies a **zombie
   zone**: real money could ride a statistically-dead edge for years inside
   the −20% kill-switch with no rule ever firing.
3. December's report has no economics-per-deployed-dollar line, inviting a
   post-hoc "the strategy works, deployment is just small" narrative — a
   reading the red team **struck** as pre-drafted excuse (§4c below).

---

## 1. (a) The scale-up ladder — PROPOSED values

Scope: applies from first real capital (a December PASS) onward. All
fractions are of the **intended allocation** — the dollar figure the owner
sets at go-live (itself sign-off item #1; the ladder governs the fraction,
not the dollars).

**Rungs (proposed): 25% → 50% → 75% → 100%.** Entry rung 25% is already
registered (`initial_deployment_fraction`). One review per rung, at
**+63 live trading days** on the current rung (the cadence the existing
comment names; ~one quarter).

**Rung criteria — ALL must hold over the trailing 63-td rung window:**

| # | Criterion | Proposed value | Source / rationale |
|---|---|---|---|
| L1 | Endogenous safety halts (memo §4 taxonomy) | **0** | Carried from `max_recent_safety_halts: 0`; ambiguity counts against us. ≥3 exogenous halts in the window is itself an operational-reliability failure (memo §4, already registered) → rung fails too. |
| L2 | Real-fill TCA median \|slippage\| | **≤ 50 bps** | The `tca_median_slippage_bps_max` criterion carried onto REAL fills — the number it was always meant for (paper fills print at NBBO and cannot measure it). |
| L3 | Live-vs-paper tracking error | **≤ 2.0% annualized** | TE = std(daily live-minus-paper return) × √252 over the rung window. Derivation of the band: at the 5%/day turnover cap and the 50 bps TCA bound, fill-noise alone implies ~0.8–1.6% annualized; 2.0% tolerates fill noise and catches structural divergence (missed rebalance, wrong book, broken parity). **The band is an appetite setting — sign-off item #4.** |
| L4 | oos-monitor DEGRADED verdicts dated in-window | **0** | One DEGRADED fails the rung; two consecutive retire outright (§5 precedence). |

**Mechanical action (proposed — the red team requires the fraction pinned):**

- **PASS all four → step UP exactly one rung.** No skipping.
- **FAIL any → step DOWN exactly one rung** (not hold). Rationale: at
  real-money scale a control failure should *shed* exposure mechanically;
  holding preserves the exposure that just failed its own checks —
  asymmetric to the house fail-closed posture. *The alternative (HOLD at
  the current rung) is listed for the owner's choice — sign-off item #5
  selects one.*
- **Floor: 25%.** A failure at the 25% rung HOLDS at 25% — the ladder
  never retires; retirement belongs exclusively to the DEGRADED rule and
  the §2 sunset.
- After a step-down, a full 63-td rung window must PASS before re-ascent.
- A rung window with an UNEVALUABLE criterion (dead comparator, §4) is a
  **HOLD** — never a step up on missing evidence, and a step down only if
  an evaluable criterion independently fails.

## 2. (b) The sunset — PROPOSED values, power stated in print

**Rule (proposed):** at the first monthly evaluation at or after
**378 clean-clock live trading days** (~18 months of 21-td months from
`oos_monitor.since` 2026-06-11 → lands **~2027-12**), if ALL of:

- S1. the go-live gate has **never passed** at any evaluation to date, AND
- S2. the placebo-ensemble percentile is **< 60** (the pinned raw
  percentile — the binding one; C5's beta-hedged companion never gates), AND
- S3. **active Sharpe ≤ 0** over the full clean-clock window (active = live
  daily return − SPY daily return, the memo §2 definition — dodging the
  beta confound),

**then RETIRE**: paper lane stops, strategy archived, the never-relitigate
covenant applies. This closes the zombie zone: a book that after 18 months
has cleared no gate, beats fewer than 60% of random books, and has earned
nothing over its own benchmark, no longer gets to consume the shop's
scarcest resource (one operator's attention) by merely not being
catastrophic.

**The power table statement — published and accepted in print (red-team
condition: this is signed, not buried):**

> At 378 td the standard error of an annualized Sharpe is
> √(252/378) ≈ 0.82. If the strategy's TRUE active Sharpe is +0.3 —
> a real but modest edge — then P(observed active Sharpe ≤ 0) =
> Φ(−0.3/0.82) ≈ **36%**. **The sunset's active-Sharpe leg alone fires on
> a true-0.3 strategy roughly one time in three.** The joint rule fires
> less often (S1–S3 must all agree), but the owner accepts, in print, that
> this sunset can retire a true-but-modest edge with material probability.
> That acceptance is the price of closing the zombie zone; signing it is
> sign-off item #8. (Conversely: a true-zero strategy *escapes* the
> active-Sharpe leg half the time — the sunset is a slow judge in both
> directions; it is a backstop, not a verdict.)

## 3. (c) The economics companion — growth per unit of exposure deployed

**Metric (report-only, printed at every rung review and every gate
evaluation):**

```
GPUE(W) = annualized live total return over window W
          ÷ mean daily gross exposure over W        (gross = Σ|position mv| / equity)
```

both from the live ledgers; alongside it, always: the mean gross itself,
live CAGR, and SPY CAGR over W. If C1's gross-cap amendment is adopted, the
denominator tracks **whichever go-live diff is currently registered** —
never a dead counterfactual (A2's red-team condition, mirrored).

**SYMMETRIC interpretation — mandated by the red team, both directions
carrying both readings:**

| Observation | Favorable reading | Equally-available unfavorable/uninformative reading |
|---|---|---|
| GPUE high / rising | the book earns efficiently per deployed dollar | a small denominator amplifying noise; a beta tape lifting a ~0.5-beta book; too few days to mean anything |
| GPUE low / falling | the edge is absent per deployed dollar | a risk-off tape dragging a beta book; deployment constraints binding (Kelly/vol-target), which is the sizing layer working as validated |

**The STRUCK reading, named:** "the strategy works — deployment is just
small" may **not** be asserted from this number (in either direction). It
was struck because this draft is written knowing the live exposure and the
modal December outcome — a privileged escape hatch drafted in advance is
not analysis. GPUE never upgrades a NOT-PASS, never downgrades a PASS,
never feeds any verdict. It exists so the December reader cannot un-know
the deployment decomposition.

**Disclosure of what was observed at drafting (2026-08-01) — the red-team
condition:** this draft was written with the frozen panel report in hand
and nothing else; no fresh live read was taken for it. Known at drafting,
from the report: live gross exposure ≈ 0.419 (= Kelly 0.140 × vol-target
scalar 2.996, report §4.1/§4.6); live Sharpe −2.34 at n=34, SE ≈ 2.7 —
statistically nothing (§4.2); the negative control +4.5% since 07-13,
beating the incumbent (the noise-floor demo working, §4.2); December's
modal outcome NOT-PASS (interpretation memo). **GPUE itself was NOT
computed at drafting.** Anyone auditing this draft later judges its
symmetry against exactly this disclosure.

## 4. The comparator assumption — named, with its failure mode

**Named assumption: the paper lane keeps trading through the real-money
era.** It is the comparator for L3 (tracking error) and the context for
the companion; the live-vs-paper decomposition is the only instrument that
separates "the strategy changed" from "execution changed."

**If the paper lane dies** (cron failure, broker paper-API sunset, account
loss) — proposed handling:

- Gaps **≤ 5 consecutive trading days**: tolerated; the dead days are
  excluded from the TE computation. **No backfill ever** — paper NBBO
  fills cannot be honestly reconstructed after the fact.
- Gaps **> 5 td in a rung window**: L3 is **UNEVALUABLE** → the rung
  review result is **HOLD** (§1 rule: no step up on missing evidence; no
  step down unless an evaluable criterion independently fails). The next
  rung window starts when the paper lane resumes.
- **Permanent death** (e.g. the vendor retires paper trading): L3 cannot
  be silently dropped — replacing or removing a ladder criterion requires
  a new owner-signed, dated registration. Until one exists, reviews HOLD.
  GPUE and the sunset are unaffected (neither needs the paper lane; S2's
  placebo replays the engine on owned data).

## 5. Precedence — pre-agreed, and never a reason to loosen DEGRADED

**Retirement via two consecutive DEGRADED verdicts supersedes everything
here.** If it fires first, (a) and (c) are moot — pre-agreed now. The
existence of a slower, gentler sunset is **never** an argument to soften,
re-time, or re-interpret the DEGRADED rule (the report's exact caution).
Order of authority when rules collide: fleet/sleeve halts → DEGRADED×2
retirement → §2 sunset → ladder. The ladder is the least-privileged rule
in the stack; it only ever sizes a book the senior rules have allowed to
exist.

## 6. What this draft deliberately does not do

- No change to any `go_live:` value, any keystone, or any live config —
  drafting is not registering.
- No new verdict instrument: L1–L4, S1–S3 are all computed from ledgers
  and monitors that already exist (`safety` halt log + memo taxonomy, TCA
  log, paper/live equity JSONLs, `oos_verdict_history.jsonl`,
  `placebo-ensemble`, `oos-monitor` benchmark block).
- No implementation: wiring the rung evaluation into `go-live-gate`'s
  report happens after sign-off, as its own change with its own tests.

## 7. Sign-off block — the owner confirms EXACTLY these numbers

Registration = every line initialed (or amended in the owner's hand), then
committed as a dated keystone-pinned `go_live:` amendment + hash-frozen
memo. An un-initialed line is not registered; an amended value is the
owner's number, which is the point (report §6.2).

| # | Number / choice | Proposed | Owner initials + date |
|---|---|---|---|
| 1 | Intended allocation (the dollars the fractions apply to) | ______ (owner-only; no proposal) | ______ |
| 2 | Rung fractions | 25% → 50% → 75% → 100% | ______ |
| 3 | Rung review cadence | 63 live td per rung | ______ |
| 4 | L3 tracking-error band (and its definition above) | 2.0% annualized | ______ |
| 5 | Failed-rung action | STEP DOWN one rung (alternative: HOLD) — circle one | ______ |
| 6 | Rung floor / ladder-never-retires | hold at 25%; retirement only via DEGRADED×2 or sunset | ______ |
| 7 | Sunset horizon | 378 clean-clock td (~2027-12) | ______ |
| 8 | Sunset legs + the accepted power number | S1 never-passed ∧ S2 placebo < 60th ∧ S3 active Sharpe ≤ 0; **accepting ~36% false-retire on a true-0.3 edge (S3 alone)** | ______ |
| 9 | Paper-lane gap tolerance / unevaluable→HOLD | 5 consecutive td | ______ |
| 10 | GPUE rendering points + symmetric-reading mandate (§3 verbatim) | every rung review + every gate evaluation | ______ |

Unchanged and not up for signature here: `max_recent_safety_halts: 0`,
TCA ≤ 50 bps (L2 restates the registered value onto real fills),
`retire_after_consecutive_degraded: 2`, the memo's halt taxonomy, and the
2026-12-01 / 126-td / 25% gate itself.

---

## (d) APPENDED 2026-08-02 — exact-money boundary + value-level reconciliation (pre-first-real-dollar)

Source: external design review 2026-08-01 (finding 3, its verified kernel —
archived with corrections at `research/2026-08-01_external_design_review.md`)
+ the standing TECH_DEBT M3 limitation. The review's urgency claim was
corrected (reconcile compares order-ID sets and equity DATES, never monetary
values, so float drift cannot trip anything today — and reconcile alerts, it
does not HALT), but the prospective kernel is right: real dollars deserve
exact arithmetic at the money boundary, and M3 says value-level
reconciliation should exist anyway.

PROPOSED (owner sign-off with the rest of this draft): before the first real
dollar, as ONE designed change with parity tests —

1. Execution-layer money arithmetic (order sizing, spend-against-cap, P&L
   comparison against broker truth) moves to exact decimal (integer cents or
   `Decimal`). The backtest engine stays float/NumPy by design — vectorized
   research arithmetic is the correct tool there; the boundary is where
   broker-truth dollars enter.
2. `reconcile` gains value-level position/quantity/equity comparison with
   explicit tolerance bands (closing TECH_DEBT M3), so divergence detection
   stops depending on ID/date presence alone.

Sequencing note: this lands BEFORE the go-live flip and AFTER December's
verdict is read (a live-path arithmetic change during the evaluation window
would churn the thing being measured).
