Thales
← research journal
Aug 1, 2026raw markdown ↗

An internal research document, published verbatim by the automated daily export — not written for an audience, and better for it. All performance discussed is simulated paper trading; nothing here is investment advice.

December governance package — DRAFT (2026-08-01)

⚠️ DRAFT — NOT REGISTERED — EVERY NUMBER REQUIRES OWNER SIGN-OFF

Panel item C6 (research/2026-08-01_quant_panel_suggestions.md §3 C6, §6.2): the ladder band, the sunset's retirement-probability acceptance, and the step-down rule are appetite settings the panel cannot make. Nothing here binds anything. This document is deliberately not hash-frozen — freezing a draft would fake a registration. Registration happens only via the §7 sign-off block: each initialed number then lands as a dated, keystone-pinned amendment to config/settings.yaml go_live: plus a hash-frozen memo (with the config_guard.py re-pin in the same commit, per house rule). Until then, the only registered December rules remain the 2026-07-05 go_live: block and the 2026-07-13 interpretation memo.

Why now: the window is low-power and December's Sharpe outcome is effectively known (report §4.2: live Sharpe −2.34 at n=34 is statistically nothing, and passing needs ~+4.0 annualized over the remaining days) — nothing below can be tuned to flatter. That is precisely what makes this the cheapest moment to write it (report: "closable now precisely because the window is low-power").

The three holes being closed (all at the capital layer, none touched by existing registrations):

  1. go_live.initial_deployment_fraction: 0.25 says "review at +63 live td before scaling further" — and names no criteria. A criteria-free review is result-shopping deferred.
  2. Between the retire rule (two consecutive DEGRADED ≈ live Sharpe −2 sustained) and a gate pass (observed ≈ +2.3 at T=126) lies a zombie zone: real money could ride a statistically-dead edge for years inside the −20% kill-switch with no rule ever firing.
  3. December's report has no economics-per-deployed-dollar line, inviting a post-hoc "the strategy works, deployment is just small" narrative — a reading the red team struck as pre-drafted excuse (§4c below).

1. (a) The scale-up ladder — PROPOSED values

Scope: applies from first real capital (a December PASS) onward. All fractions are of the intended allocation — the dollar figure the owner sets at go-live (itself sign-off item #1; the ladder governs the fraction, not the dollars).

Rungs (proposed): 25% → 50% → 75% → 100%. Entry rung 25% is already registered (initial_deployment_fraction). One review per rung, at +63 live trading days on the current rung (the cadence the existing comment names; ~one quarter).

Rung criteria — ALL must hold over the trailing 63-td rung window:

#CriterionProposed valueSource / rationale
L1Endogenous safety halts (memo §4 taxonomy)0Carried from max_recent_safety_halts: 0; ambiguity counts against us. ≥3 exogenous halts in the window is itself an operational-reliability failure (memo §4, already registered) → rung fails too.
L2Real-fill TCA median |slippage|≤ 50 bpsThe tca_median_slippage_bps_max criterion carried onto REAL fills — the number it was always meant for (paper fills print at NBBO and cannot measure it).
L3Live-vs-paper tracking error≤ 2.0% annualizedTE = std(daily live-minus-paper return) × √252 over the rung window. Derivation of the band: at the 5%/day turnover cap and the 50 bps TCA bound, fill-noise alone implies ~0.8–1.6% annualized; 2.0% tolerates fill noise and catches structural divergence (missed rebalance, wrong book, broken parity). The band is an appetite setting — sign-off item #4.
L4oos-monitor DEGRADED verdicts dated in-window0One DEGRADED fails the rung; two consecutive retire outright (§5 precedence).

Mechanical action (proposed — the red team requires the fraction pinned):

  • PASS all four → step UP exactly one rung. No skipping.
  • FAIL any → step DOWN exactly one rung (not hold). Rationale: at real-money scale a control failure should shed exposure mechanically; holding preserves the exposure that just failed its own checks — asymmetric to the house fail-closed posture. The alternative (HOLD at the current rung) is listed for the owner's choice — sign-off item #5 selects one.
  • Floor: 25%. A failure at the 25% rung HOLDS at 25% — the ladder never retires; retirement belongs exclusively to the DEGRADED rule and the §2 sunset.
  • After a step-down, a full 63-td rung window must PASS before re-ascent.
  • A rung window with an UNEVALUABLE criterion (dead comparator, §4) is a HOLD — never a step up on missing evidence, and a step down only if an evaluable criterion independently fails.

2. (b) The sunset — PROPOSED values, power stated in print

Rule (proposed): at the first monthly evaluation at or after 378 clean-clock live trading days (~18 months of 21-td months from oos_monitor.since 2026-06-11 → lands ~2027-12), if ALL of:

  • S1. the go-live gate has never passed at any evaluation to date, AND
  • S2. the placebo-ensemble percentile is < 60 (the pinned raw percentile — the binding one; C5's beta-hedged companion never gates), AND
  • S3. active Sharpe ≤ 0 over the full clean-clock window (active = live daily return − SPY daily return, the memo §2 definition — dodging the beta confound),

then RETIRE: paper lane stops, strategy archived, the never-relitigate covenant applies. This closes the zombie zone: a book that after 18 months has cleared no gate, beats fewer than 60% of random books, and has earned nothing over its own benchmark, no longer gets to consume the shop's scarcest resource (one operator's attention) by merely not being catastrophic.

The power table statement — published and accepted in print (red-team condition: this is signed, not buried):

At 378 td the standard error of an annualized Sharpe is √(252/378) ≈ 0.82. If the strategy's TRUE active Sharpe is +0.3 — a real but modest edge — then P(observed active Sharpe ≤ 0) = Φ(−0.3/0.82) ≈ 36%. The sunset's active-Sharpe leg alone fires on a true-0.3 strategy roughly one time in three. The joint rule fires less often (S1–S3 must all agree), but the owner accepts, in print, that this sunset can retire a true-but-modest edge with material probability. That acceptance is the price of closing the zombie zone; signing it is sign-off item #8. (Conversely: a true-zero strategy escapes the active-Sharpe leg half the time — the sunset is a slow judge in both directions; it is a backstop, not a verdict.)

3. (c) The economics companion — growth per unit of exposure deployed

Metric (report-only, printed at every rung review and every gate evaluation):

GPUE(W) = annualized live total return over window W
          ÷ mean daily gross exposure over W        (gross = Σ|position mv| / equity)

both from the live ledgers; alongside it, always: the mean gross itself, live CAGR, and SPY CAGR over W. If C1's gross-cap amendment is adopted, the denominator tracks whichever go-live diff is currently registered — never a dead counterfactual (A2's red-team condition, mirrored).

SYMMETRIC interpretation — mandated by the red team, both directions carrying both readings:

ObservationFavorable readingEqually-available unfavorable/uninformative reading
GPUE high / risingthe book earns efficiently per deployed dollara small denominator amplifying noise; a beta tape lifting a ~0.5-beta book; too few days to mean anything
GPUE low / fallingthe edge is absent per deployed dollara risk-off tape dragging a beta book; deployment constraints binding (Kelly/vol-target), which is the sizing layer working as validated

The STRUCK reading, named: "the strategy works — deployment is just small" may not be asserted from this number (in either direction). It was struck because this draft is written knowing the live exposure and the modal December outcome — a privileged escape hatch drafted in advance is not analysis. GPUE never upgrades a NOT-PASS, never downgrades a PASS, never feeds any verdict. It exists so the December reader cannot un-know the deployment decomposition.

Disclosure of what was observed at drafting (2026-08-01) — the red-team condition: this draft was written with the frozen panel report in hand and nothing else; no fresh live read was taken for it. Known at drafting, from the report: live gross exposure ≈ 0.419 (= Kelly 0.140 × vol-target scalar 2.996, report §4.1/§4.6); live Sharpe −2.34 at n=34, SE ≈ 2.7 — statistically nothing (§4.2); the negative control +4.5% since 07-13, beating the incumbent (the noise-floor demo working, §4.2); December's modal outcome NOT-PASS (interpretation memo). GPUE itself was NOT computed at drafting. Anyone auditing this draft later judges its symmetry against exactly this disclosure.

4. The comparator assumption — named, with its failure mode

Named assumption: the paper lane keeps trading through the real-money era. It is the comparator for L3 (tracking error) and the context for the companion; the live-vs-paper decomposition is the only instrument that separates "the strategy changed" from "execution changed."

If the paper lane dies (cron failure, broker paper-API sunset, account loss) — proposed handling:

  • Gaps ≤ 5 consecutive trading days: tolerated; the dead days are excluded from the TE computation. No backfill ever — paper NBBO fills cannot be honestly reconstructed after the fact.
  • Gaps > 5 td in a rung window: L3 is UNEVALUABLE → the rung review result is HOLD (§1 rule: no step up on missing evidence; no step down unless an evaluable criterion independently fails). The next rung window starts when the paper lane resumes.
  • Permanent death (e.g. the vendor retires paper trading): L3 cannot be silently dropped — replacing or removing a ladder criterion requires a new owner-signed, dated registration. Until one exists, reviews HOLD. GPUE and the sunset are unaffected (neither needs the paper lane; S2's placebo replays the engine on owned data).

5. Precedence — pre-agreed, and never a reason to loosen DEGRADED

Retirement via two consecutive DEGRADED verdicts supersedes everything here. If it fires first, (a) and (c) are moot — pre-agreed now. The existence of a slower, gentler sunset is never an argument to soften, re-time, or re-interpret the DEGRADED rule (the report's exact caution). Order of authority when rules collide: fleet/sleeve halts → DEGRADED×2 retirement → §2 sunset → ladder. The ladder is the least-privileged rule in the stack; it only ever sizes a book the senior rules have allowed to exist.

6. What this draft deliberately does not do

  • No change to any go_live: value, any keystone, or any live config — drafting is not registering.
  • No new verdict instrument: L1–L4, S1–S3 are all computed from ledgers and monitors that already exist (safety halt log + memo taxonomy, TCA log, paper/live equity JSONLs, oos_verdict_history.jsonl, placebo-ensemble, oos-monitor benchmark block).
  • No implementation: wiring the rung evaluation into go-live-gate's report happens after sign-off, as its own change with its own tests.

7. Sign-off block — the owner confirms EXACTLY these numbers

Registration = every line initialed (or amended in the owner's hand), then committed as a dated keystone-pinned go_live: amendment + hash-frozen memo. An un-initialed line is not registered; an amended value is the owner's number, which is the point (report §6.2).

#Number / choiceProposedOwner initials + date
1Intended allocation (the dollars the fractions apply to)______ (owner-only; no proposal)______
2Rung fractions25% → 50% → 75% → 100%______
3Rung review cadence63 live td per rung______
4L3 tracking-error band (and its definition above)2.0% annualized______
5Failed-rung actionSTEP DOWN one rung (alternative: HOLD) — circle one______
6Rung floor / ladder-never-retireshold at 25%; retirement only via DEGRADED×2 or sunset______
7Sunset horizon378 clean-clock td (~2027-12)______
8Sunset legs + the accepted power numberS1 never-passed ∧ S2 placebo < 60th ∧ S3 active Sharpe ≤ 0; accepting ~36% false-retire on a true-0.3 edge (S3 alone)______
9Paper-lane gap tolerance / unevaluable→HOLD5 consecutive td______
10GPUE rendering points + symmetric-reading mandate (§3 verbatim)every rung review + every gate evaluation______

Unchanged and not up for signature here: max_recent_safety_halts: 0, TCA ≤ 50 bps (L2 restates the registered value onto real fills), retire_after_consecutive_degraded: 2, the memo's halt taxonomy, and the 2026-12-01 / 126-td / 25% gate itself.


(d) APPENDED 2026-08-02 — exact-money boundary + value-level reconciliation (pre-first-real-dollar)

Source: external design review 2026-08-01 (finding 3, its verified kernel — archived with corrections at research/2026-08-01_external_design_review.md)

  • the standing TECH_DEBT M3 limitation. The review's urgency claim was corrected (reconcile compares order-ID sets and equity DATES, never monetary values, so float drift cannot trip anything today — and reconcile alerts, it does not HALT), but the prospective kernel is right: real dollars deserve exact arithmetic at the money boundary, and M3 says value-level reconciliation should exist anyway.

PROPOSED (owner sign-off with the rest of this draft): before the first real dollar, as ONE designed change with parity tests —

  1. Execution-layer money arithmetic (order sizing, spend-against-cap, P&L comparison against broker truth) moves to exact decimal (integer cents or Decimal). The backtest engine stays float/NumPy by design — vectorized research arithmetic is the correct tool there; the boundary is where broker-truth dollars enter.
  2. reconcile gains value-level position/quantity/equity comparison with explicit tolerance bands (closing TECH_DEBT M3), so divergence detection stops depending on ID/date presence alone.

Sequencing note: this lands BEFORE the go-live flip and AFTER December's verdict is read (a live-path arithmetic change during the evaluation window would churn the thing being measured).