Thales
← research journal
Jul 13, 2026raw markdown ↗

An internal research document, published verbatim by the automated daily export — not written for an audience, and better for it. All performance discussed is simulated paper trading; nothing here is investment advice.

December Gate — Interpretation Memo (pre-registered 2026-07-13)

Status: BINDING INTERPRETATION, append-only after today. Written while the clean forward window contains zero safety halts, zero gate evaluations, and ~20 of 126 trading days — i.e., before any fact exists that could motivate the interpretations below. The pinned gate criteria in config/settings.yaml go_live: are NOT changed by this memo (changing them mid-window is gate-shopping); this memo pre-commits how their outputs will be read. External critique 2026-07-13 prompted it; the arithmetic was verified independently before adoption.


1. The power of the Sharpe criterion, stated plainly

The binding criterion min_forward_sharpe_lower: 0.0 (bootstrap 95% lower bound of live Sharpe > 0) has, at the 126-day minimum window:

  • Standard error of an annualized Sharpe ≈ √(252/T). At T=126: ≈ 1.41.
  • To clear a 95% lower bound > 0, observed live Sharpe must be ≈ 2.3+.
  • If the honest true Sharpe is ~0.65 (the survivorship-free read), P(pass at 126d) ≈ 12%. If the true Sharpe is 0, P(pass) = 5% by construction.

Pre-committed reading: a December NOT-PASSED on the Sharpe criterion is the modal outcome under the honest prior and is weak evidence about edge — it must not be read as "the strategy failed." A December PASS at T≈126 is mostly luck even when real edge exists; the deployment fraction (25%) and paper haircut (0.5×) already price that in. The window keeps growing after December: at T=252 the required observed Sharpe falls to ~1.6, at T=504 to ~1.2. Proving Sharpe ≈ 0.6 statistically takes 3–5 years; the gate's real December job is §3, not alpha proof.

2. Beta contamination and the companion metric

The binding Sharpe is total-return. A long-only, vol-targeted equity book is substantially a market bet: a melt-up half-year can pass on beta with zero alpha; a flat tape can fail genuine alpha. Pre-committed remedy (report-only, implemented 2026-07-13 in oos-monitorbenchmark block):

  • Active Sharpe = Sharpe of (live daily return − SPY daily return), plus realized beta and correlation over the live window.
  • Reading: a PASS with active Sharpe ≤ 0 is a beta pass — deploy-eligible per the letter of the gate, but the memo pre-commits that the December decision notes say so explicitly, and the initial deployment stays at the 25% floor (no discretionary upsizing on a beta pass). A NOT-PASS with clearly positive active Sharpe in a down/flat tape is recorded as "alpha-consistent, beta-failed" — grounds to continue paper, never to override the gate.
  • The companion NEVER feeds any verdict. It exists so the human reading December's report cannot un-know the decomposition.

3. What December actually decides

An operational-competence gate authorizing a small, capped experiment — not statistical proof of alpha. 126 days CAN measure with high power: clean execution (zero endogenous halts), costs (TCA median ≤ 50 bps), state integrity (reconcile clean, no manual surgeries), engine-live parity, and process discipline (no gate-shopping, ledgers intact). It CANNOT measure Sharpe ≈ 0.6 vs 0. The pinned criteria already embody this (haircut 0.5×, initial deployment 25%, manual checklist); this section makes it explicit so nobody — including future-us — retells December as "the strategy proved itself."

4. Halt taxonomy (pre-registered before any halt exists)

The binding count (max_recent_safety_halts: 0) counts all halts — that pin is untouched. Interpretation, pre-committed now and mirrored in execution/go_live.py:classify_halt:

  • ENDOGENOUS (strategy/system fault — argues against go-live): account identity mismatch; daily-loss breach; runaway order count; total-notional breach; option-structure violations; any unknown/unclassified reason (ambiguity counts against us, never for us).
  • EXOGENOUS (availability event — must not fail the gate on a technicality): broker unreachable; broker account/trading blocked by vendor; degenerate vendor equity; operator-engaged manual/fleet halts (the operator halting out of caution is prudence, not a strategy fault — the halt's logged reason governs if it references a strategy fault).
  • Pre-committed reading: ≥1 endogenous halt → the criterion fails on the merits. Exogenous-only halts → the criterion technically fails (pinned count is binding) but the December decision notes record it as an availability event; ≥3 exogenous halts in the window is itself an operational-reliability failure (the "competence" this gate measures includes tolerating vendor weather).

5. The economic bar (pre-registered formula; rates to be confirmed)

A statistical pass must also clear economics before real money: after-tax live CAGR > after-tax SPY buy-and-hold CAGR over the same window, where the strategy's gains are taxed at the operator's short-term marginal rate (monthly turnover ⇒ predominantly short-term) and B&H at the long-term rate. Placeholder rates ST 32% / LT 15% — the operator confirms actual brackets before 2026-12-01; the formula is pinned now so the rates can't be chosen after seeing which choice flatters the result. At current account scale the dollars are small; the discipline is the point.

6. VRP forward-gate context (regime, fills)

  • One window = one regime draw. At VRP's gate evaluation, record: the window's mean VIX percentile vs 2004+ history, and whether a ≥5% SPY drawdown occurred in-window. Pre-committed discount: a PASS earned in a bottom-tercile-VIX, no-drawdown window is labeled "calm-regime pass — unstressed" and does not by itself justify scaling; the stressed behavior is the thing the premium is paid for.
  • Paper option fills are optimistic (NBBO, no queue/spread pain). The pessimistic-twin ledger (implemented 2026-07-13, data/state/vrp/pessimistic_ledger.jsonl) records every open/close at the worst side of the same-moment NBBO. Pre-committed reading: VRP's honest P&L lives between the paper curve and the twin curve; the gate evaluation reports both, and a pass that exists only on the optimistic curve is not a pass.

7. The placebo and its limits

meanrev (the negative control) is n=1 — one draw from the dead-strategy noise distribution. It validates infrastructure and provides a visceral benchmark, not calibration. Planned (deferred, design pinned): a placebo ensemble — ≥100 random 50-name selections replayed through the engine over the same forward window as momentum's clean clock, equal-weighted, costed — producing a null Sharpe distribution; momentum's live Sharpe is then reported as a percentile of that ensemble. Build when the forward window is long enough to make it meaningful (~60+ days); its design is recorded here first so the ensemble's parameters can't be tuned to flatter the incumbent.

8. Change discipline

This memo is append-only. Any future edit adds a dated section; nothing above this line is rewritten after 2026-07-13. The taxonomy marker list in go_live.py may gain NEW exogenous markers only for halt reasons that do not yet exist in the codebase (a new vendor failure mode), never reclassify an existing one after it has fired. Git history is the timestamp; external timestamping (OpenTimestamps) is offered but not required while the audience is the operator.


9. Appended 2026-07-19 (does not alter anything above the §8 freeze line)

Economic-bar rates confirmed (operator's federal marginal bracket). The §5 placeholder (ST 32% / LT 15%) is replaced for evaluation by the operator's actual rates: short-term 24% (ordinary income — the strategy's monthly-turnover gains) and long-term 15% (LTCG — the SPY buy-and-hold benchmark). State income tax: none declared as of this append; if applicable it is added to BOTH rates (it only widens the hurdle, the honest direction). The §5 FORMULA is unchanged — only the rates are filled in, exactly as §5 anticipated. So "pass" = strategy after-24%-tax CAGR > SPY-B&H after-15%-tax CAGR over the window, on top of the statistical gate.

Placebo-ensemble parameters PINNED (the §7 build, backtest/placebo.py). Built 2026-07-19 while the forward window is still low-power (22 days), so the harness's parameters are fixed BEFORE they judge decision-grade data:

  • n_placebos 200, sleeve_size 50, cost 10 bps round-trip, rebalance monthly, seed_base 20260713 (fixed — reproducible, not cherry-picked), min_days_for_verdict 60, Sharpe annualized 252.
  • Null = random equal-weight 50-name portfolios drawn from the PiT-captured investable universe over the SAME forward window; target = the sleeve's live daily returns. Report-only companion to oos-monitor; it NEVER gates or trades.
  • Pinned interpretation: live at ≥90th percentile of the null = signal beyond random stock-picking; ≤60th = indistinguishable from luck; between = weak/ambiguous. Below 60 live days the percentile is NOT decision-grade (stated on every run).
  • Consistency check at build: the null's median Sharpe matched SPY's independently-computed Sharpe over the same window (random baskets ≈ market), and the target matched the benchmark-companion excess — the harness is internally consistent.
  • tests/test_backtest/test_placebo.py asserts the code matches these pins (lockstep). Any change to a pinned value re-defines the null and is a further dated append here — never a silent edit.

10. Appended 2026-08-01 (does not alter anything above the §8 freeze line)

Beta-hedged companion for the placebo ensemble — REGISTERED (panel item C5, 2026-08-01 quant panel, validation methodologist; red-team revisions folded in).

The §9-pinned null carries full market exposure (beta ≈ 1: random 50-name equal-weight books ARE the market) while the live book runs ~0.49 gross. In an up-tape the raw percentile therefore handicaps the incumbent for reasons unrelated to stock picking; in a down-tape it flatters — exactly as the window approaches decision-grade (~60 days, expected late September). This section registers a REPORT-ONLY companion NOW, before the null is decision-grade, so neither its design nor its interpretation can be tuned to a known outcome.

Method (implemented in backtest/placebo.py:beta_hedged_companion, rendered by thales placebo-ensemble): estimate each book's beta to SPY over the same window (OLS slope of the book's daily returns on SPY daily returns); subtract beta × SPY_return from EVERY null book's daily returns AND from the live curve; recompute Sharpes; report the hedged live Sharpe's percentile of the hedged null NEXT TO the raw pinned percentile. The raw §9-pinned percentile remains the ONLY binding number — changing the pinned null mid-window would be the actual sin; the companion never gates, never trades, and is labeled non-binding on every render.

Observed at registration (2026-08-01 — recorded so pre-dating is auditable):

  • Clean-clock live window: 34 trading days since 2026-06-11 (committed equity_history.jsonl); NOT decision-grade under the §9 pin (min 60).
  • Raw pinned read today: live Sharpe −2.344, 0th percentile of the 200-book null (null median +2.469, p5 +0.017, p95 +4.306).
  • Tape direction over the window: UP — the PiT-captured panel's equal-weight cumulative return is +3.49% (2026-06-11 → 2026-07-30, owned forward capture), and the entire null Sharpe distribution sitting at/above ~0 says the same thing (random baskets ≈ market); consistent with the July digest record.
  • Therefore the companion TODAY flatters the incumbent: hedging removes the market tailwind the beta≈1 nulls enjoy and the beta≈0.49 live book lacks, so today's hedged percentile can only sit at-or-above the raw one. Registering in this state is deliberate and disclosed; the correction is symmetric and will penalize the incumbent in a down-tape.
  • The companion has NOT been computed on live data as of this append: the owned SPY daily series (data/raw/SPY.parquet) ends 2026-06-05 — before the clean-clock window — so no hedged percentile existed to peek at. The companion renders once SPY closes covering the window are on disk; until then thales placebo-ensemble prints the raw number plus an explicit companion-unavailable note.

Pre-committed NEUTRALIZED divergence reading (binding interpretation; the original privileged reading was struck by the red team): a raw-vs-hedged divergence means ONLY that the beta confound is material over this window. Neither percentile is privileged. Divergence may not be read as "the hedged number exonerates the incumbent" nor as "the raw number stands regardless" — it is a measurement about the instrument, not about skill.

Beta-estimation caveat (carried on EVERY render): on windows this short the OLS beta's standard error is ~0.15–0.25, so the hedged percentile inherits material estimation noise; single-render divergences are weak evidence.

Kill criterion (pinned): if by 2027-02-01 (6 months from registration) the maximum absolute raw-vs-hedged divergence observed across decision-grade renders is < 15 percentile points, the companion is INERT — retire it and record the null here. No other retirement or promotion path exists.

Lockstep tests: tests/test_backtest/test_placebo.py (companion math on fixed-beta fixtures; render carries both percentiles + the caveat + the raw-is-binding line). The §9 raw computation is asserted byte-identical through the refactor that exposes per-book return paths.