Apparatus Readiness — 2026-05-30 AFK session index
Theme: NON-alpha. Harden the apparatus that will judge and execute the edge
before real money. 0 SHIPs by design — success = trustworthy tooling. Backtest
byte-identical throughout (sharpe 0.9321437890531372, 2010-01-01→2026-03-17).
This is the reviewer-facing index; detailed notes are linked.
What now exists (three faces of the apparatus)
1. Judging — forward OOS degradation monitor
- Code:
backtest/oos_monitor.py; CLI:thales oos-monitor [--since DATE] [--power] [--json-out P]. - Grades LIVE returns vs the in-sample backtest by a pre-registered rule (thresholds fixed in-module before looking at live data).
- Verdicts:
INSUFFICIENT_DATA(N < 60) /OK/DEGRADED. - Wired into audit:
AUDIT.mdM5 (monthly) + a ~2026-08 first-verdict one-off. Monthly-revalidation.yml step is documented for one-time human paste (AFK hook blocks.github/workflowsedits). - Detail:
research/2026-05-30_oos_monitor_power.md.
2. Cost realism — capacity / stress-execution study
- Code:
backtest/capacity.py; CLI:thales capacity-report. - Shortfall by AUM tier (normal rebalance + kill-switch liquidation, calm & stressed) via the EXISTING Almgren-Chriss model (reused, not rewritten).
- Wired into audit: event-driven AUM trigger in
AUDIT.md. - Detail:
research/2026-05-30_capacity_stress.md.
3. Execution safety — kill-switch residual-position invariant
execution/health.py:evaluate_account_health(..., kill_switch_active=)hard-fails if the book is risk-off but positions remain (failed liquidation). Wired into the daily CI health gate viathales portfolio health.
Two real defects this session found and fixed (build-then-verify)
- Monitor cried wolf ~40%. The first binding rule (live Sharpe point < IS CI lower) ignored the live estimate's own sampling error → 27–43% false-alarm rate when live truly equals in-sample. Fixed: flag DEGRADED only when the live Sharpe's bootstrap UPPER bound < in-sample point → false-alarm ~3%.
- Capacity note misstated liquidation. I had claimed the turnover cap would
stage a kill-switch exit (making single-day cost an upper bound). Verified in
code it's FALSE: the kill-switch bypasses the cap (engine
new_weights={}after the cap; live_liquidate_to_cashuncapped). The whole book is dumped in one session — the single-day stressed cost is the actual behavior.
Pre-real-money fix priority (from the execution + data + validation audit, turns 14-32)
The audit of the existing code (read-only; documented in TECH_DEBT.md, not fixed — BRIEFING) found these. Fix order before real money:
- MUST FIX (P0) — CPCV validator leaks the test set into in-sample (validation
integrity).
cpcv.py:_run_single_pathruns the IS backtest on the contiguous[train_start,train_end]range, not the purged non-contiguoustrain_idx, so 6/15 default paths have IS ⊇ OOS (VERIFIED; all 6 run the identical full-sample IS → constant IS Sharpe 0.345, 40% of paths degenerate). The reported PBO/IS-OOS-corr/DSR are contaminated (corrupted in a data-dependent direction; e.g. measured IS-OOS corr 0.089 reported vs 0.152 clean-only). This is the tool that "validated" every keystone, so the whole validation rests partly on it. Confined to cpcv.py (walk-forward screen is clean). Fix is foundational (engine can't run non-contiguous dates) → re-validate everything. Pinned bytest_cpcv_no_is_oos_leakage. - MUST FIX (P0) — live vol-delever overlay is broken.
daily._check_vol_delevercrashes in production (read_equity_history()returnslist[float], code indexes dicts — VERIFIED at runtime; masked by tests that mock the wrong contract) AND sizes shares asint(notional/equity*100)(nonsense). Latent until ≥64 live days (~Aug 2026). It's live-only (not in the validated backtest) → cleanest fix may be to DELETE it. Pinned by an xfail test. The contract-mismatch bug class is bounded to this one site. - SHOULD FIX (P2):
- Data freshness: the trade path has no signal-freshness assertion;
validatestaleness is lax (10 cal-days vs 5-trading-day intent) + Monday-only. "Trades on stale signals" is under-defended (heartbeat catches fetch errors, not stale-success). - Idempotency:
already_ran_todaykeys on the last-written run_summary + noclient_order_id→ a crash-window same-day re-run with pending orders could double-submit (narrow; positions-move mitigant). - Kelly snapshot records target names incl. failed buys → phantom returns (narrow).
- Data freshness: the trade path has no signal-freshness assertion;
- NICE (P3): broker
get_orders_sincesilent 500-cap (weakens reconcile);submit_market_ordernon-"buy"→SELL default;validatethresholds.
Honest limits — what this apparatus CANNOT do
- The historical overfitting numbers are partly contaminated. Per the CPCV P0 above, 6/15 CPCV paths leak (IS ⊇ OOS), so the reported PBO / IS-OOS-corr / DSR cited across CLAUDE.md / RESEARCH.md are computed over invalid paths. Don't treat the exact CPCV figures as ground truth until the leak is fixed and CPCV re-run. (The walk-forward screen Sharpe Δ / CIs are unaffected — leak is cpcv.py-only.)
- The monitor is a catastrophe backstop, not an early-warning. Sharpe is too
noisy: it reliably flags only a severe, sustained collapse (Sharpe → ≤0 over
years; ~91% power at 5y). A true halving of the edge is ~18% detectable even at
5y. A passing
OKdoes NOT certify the edge. - No backtest closes the IS/OOS gap. The only true OOS test is the multi-month live paper A/B; this tooling judges it rigorously but cannot substitute for it.
- Capacity ≈ $400M ceiling (flat-10bps breaks >15bps), driven by a few thin-ADV names (MLI/ENS/BWA). Fine far above any near-term plan. Paper TCA (~0bps) is NOT a capacity signal (simulated NBBO fills at trivial size).
- The judge is only as good as its inputs. Guards added:
--since(avoid mixing a prior live config's returns), in-sample provenance + shortness warning (avoid a truncated reference), residual-position check (avoid silent post-crash exposure). These mitigate, not eliminate, garbage-in risk.
Pre-real-money runbook (what to run, when)
- GATING — before any real money: clear the P0 (vol-delever) and decide on the P2s in "Pre-real-money fix priority" above. The P0 self-activates ~Aug 2026 even on paper.
- Now → ongoing (daily CI):
thales portfolio health— account + book sanity, incl. the kill-switch-flat + drawdown-headroom checks;thales portfolio kill-switch-status— one-shot email on a fresh fire. Local integrity in one command:thales selfcheck(config-drift W3 + data-freshness + data-quality W5 + engine-regression W2). - ~2026-06-01: verify the live book transitioned to HRP+Kelly (AUDIT one-off).
- ~2026-08 (≥60 live days):
thales oos-monitor --since <HRP-transition date>— first non-INSUFFICIENT verdict; expect OK/INSUFFICIENT (low power by design). - Monthly: survivorship-free CPCV (monthly-revalidation.yml) + oos-monitor M5.
- If AUM → ~$100M:
thales capacity-report; switch backtest to the impact model and re-validate; decide kill-switch staging (currently single-session).
Session arc
Phase 1 (build, turns 1-13): the monitor + capacity tooling, hardened, reviewed, integration-tested, wired into the audit, and consolidated into this index. Phase 2 (audit, turns 14-32): adversarial read-only audit of the EXISTING execution, data, alerting, AND validation code → 2 P0 (live vol-delever crash; CPCV IS/OOS leakage), 4 P2, 3 P3 (TECH_DEBT.md), plus 7 shipped additive gates/commands (kill-switch-residual, data-freshness, drawdown-headroom, verify-baseline, verify-config, kill-switch fresh-fire alert, selfcheck). Every finding verified + severity-honest (agent overclaims and my own turn-30 direction guess downgraded/corrected); contract-mismatch and CPCV-leak blast radii bounded. Production byte-identical throughout; 0 SHIPs (correct for a non-alpha session). The biggest result: the tool that validated every keystone (CPCV) has a test-set leak — the apparatus-hardening mandate surfaced a flaw in the judge itself.