2026-04-06 Research Notes
17:46 — Parallel CPCV validated end-to-end on production config
Context: Commits 5aef2dd (parallel CPCV with --workers) and cfa4c34 (strategy-class pickling + worker error surfacing fix) needed a real-data sanity check. Used this as an opportunity to refresh the CPCV baseline on the current production config (momentum + smoothness + Kelly + HRP).
Command: caffeinate -i -m -s thales cpcv --workers 6
Result:
- Wall time: 3m 57s (start 17:42:55, finish 17:46:52)
- Sequential baseline estimate: ~30 min → ~7.5× speedup with 6 workers on M5 (4P + 6E cores)
- 6 workers saturated cores at 120-127% CPU each (numpy/polars thread pools on top of multiprocessing)
- Memory: ~4% each × 6 = ~7.5 GB additional, comfortable within free headroom
- Results identical to what sequential would produce: PBO 50%, 15/15 positive OOS paths
Full numbers (results/diagnostics/cpcv_2026-04-06T17-46-52.json):
| Metric | Value |
|---|---|
| PBO | 50.0% |
| Observed Sharpe | 0.611 |
| Deflated Sharpe | 0.000 |
| Mean OOS Sharpe | +0.61 (range 0.35 – 1.01) |
| Mean OOS CAGR | +5.1% (range 2.6% – 9.7%) |
| Max OOS drawdown | 44.2% (crisis paths) |
| Positive OOS paths | 15/15 |
| IS-OOS correlation | +0.02 (weak) |
Note on delta from MEMORY.md baseline: MEMORY.md had mean OOS Sharpe 0.799 and max OOS DD 1.0–3.8% from the pre-HRP ablation. Fresh run shows 0.611 and 6.4–44.2%. The widened DD range is because the fresh CPCV windows straddle the 2008 GFC (kill-switch still firing at ~22% in the log — risk controls are intact, the windows just now include crisis periods). The lower mean Sharpe is partly HRP's "crisis insurance" effect — see the next entry.
Pre-run fix (cfa4c34):
- Worker was resolving strategy class by name with a silent fallback to
MomentumStrategy— anyMultifactorStrategyor new-class run would silently get momentum results. Replaced with direct pickling ofstrategy_cls. - Worker exceptions were swallowed (returned
None), making failed paths indistinguishable from "insufficient data" skips. Now workers return(path_idx, path_or_none, error_or_none)and the main process logs vialogger.warning. - All 18 CPCV tests still pass.
18:02 — HRP ablation: walk-forward vs CPCV disagree; HRP stays on
Question: Earlier runs showed mean OOS Sharpe dropped from ~0.799 (pre-HRP baseline) → 0.611 (current HRP config). Was HRP actually helping or is it dead weight? Commit 9c1f726 claimed "HRP CPCV matches baseline" but the delta looks unfavorable.
Design: Two-test comparison.
- Walk-forward A/B with
thales evaluate -p strategy.construction.weighting=equal— 5 rolling windows, 1000-sample bootstrap CIs - Full-history CPCV with
strategy.construction.weighting: equal— same 15-path setup as HRP baseline
Walk-forward result (eval_hrp_off.json, 5 windows 2017-2026)
| Metric | Baseline (HRP) | Candidate (equal) | Δ | 95% CI | P(Δ>0) | Wins | Verdict |
|---|---|---|---|---|---|---|---|
| Sharpe | 0.910 | 1.093 | +0.181 | [+0.011, +0.367] | 98% | 4/5 | BETTER |
| CAGR | 8.45% | 13.97% | +5.13% | [+0.74%, +10.01%] | 99% | 5/5 | BETTER |
| MaxDD | 15.48% | 15.91% | +0.43% | [−10.20%, +5.80%] | 50% | 2/5 | INCONCLUSIVE |
Per window: W1 Sharpe 0.62→0.79, W2 0.31→0.44, W3 0.77→1.23 (CAGR 14.7→31.7%!), W4 1.71→1.96, W5 1.14→1.05 (only window where HRP wins). Walk-forward says equal weight is clearly better at 98–99% confidence.
CPCV result with equal weight (cpcv_2026-04-06T18-02-49.json, full 2005-2026)
| Metric | HRP on (prior run) | HRP off (this run) | Change |
|---|---|---|---|
| PBO | 50.0% | 75.0% | +25pp worse |
| Mean OOS Sharpe | +0.61 | +0.56 | −0.05 |
| OOS Sharpe range | 0.35 – 1.01 | −0.09 – 1.11 | 1 path went negative |
| Positive OOS paths | 15/15 | 14/15 | regression |
| Max OOS drawdown | 44.2% | 63.2% | +19pp worse |
| IS-OOS correlation | +0.02 | −0.91 | catastrophic |
The IS-OOS correlation of −0.91 is a screaming overfitting flag — it literally means the paths that looked best in training became the worst out-of-sample. CPCV's own report says "Negative IS-OOS correlation suggests overfitting".
Worst path (test groups 0,1): IS Sharpe +1.06 → OOS Sharpe −0.09, OOS MaxDD 63.2% — that's the 2005-2009 window eating the strategy alive.
Mechanism
Equal weight concentrates on high-vol momentum winners. In 2017-26 those are MAG7 compounders — equal weight catches the whole upside. In 2008 those were financials and dot-com survivors near peaks — the portfolio gets destroyed. HRP clusters correlated stocks and allocates inversely to cluster volatility, so it trims concentration in the high-vol regime-sensitive winners. That costs ~5% CAGR in bull regimes but saves ~20pp of max drawdown and ~25pp of PBO across crisis-inclusive windows.
Why walk-forward misled us
The harness only tests 2017-2026 — entirely post-GFC, only COVID as a "crisis" (brief V-shape). It never stressed the 2008 failure mode where equal-weight momentum blows up. CPCV with full history is specifically designed to exercise the strategy across crises and across non-contiguous train/test splits, and it's the canonical anti-overfitting test. When walk-forward and CPCV disagree, CPCV wins — especially when PBO moves and IS-OOS correlation flips sign.
Decision
Keep HRP on. Accept the ~5% CAGR cost as crisis insurance. Annotated config/settings.yaml with a comment capturing the trade-off. Commit e8dfe19 pushed.
Methodology lesson (for future experiments)
Walk-forward that only covers post-2008 is not a reliable overfitting detector. It cannot catch crisis-regime failure modes that don't appear in the test window. Default to CPCV with the full data history for any decision that affects weighting, sizing, or risk control. Walk-forward is still useful for signal-level A/B where the failure mode is the strategy under-fitting recent data, not overfitting crises.
Artifacts
results/evaluations/eval_hrp_off.json(walk-forward A/B)results/diagnostics/cpcv_2026-04-06T17-46-52.json(HRP on, baseline)results/diagnostics/cpcv_2026-04-06T18-02-49.json(HRP off)
Related commits
cfa4c34— parallel CPCV strategy-class + error-surfacing fixe8dfe19— HRP ablation documented inconfig/settings.yaml
18:30 — Research notebook initialized
Created research/ folder with per-day notes (YYYY-MM-DD.md) and a README.md holding the Current State snapshot plus daily log index. Superseded the single RESEARCH.md approach from earlier in the day after discussion.
research/README.md— comprehensive Current State (universe/data, production strategy, risk, execution, ML application status, performance snapshot, built-but-disabled components, open questions, infrastructure) + daily log indexresearch/2026-04-06.md— this file- Updated
memory/feedback_research_notebook.mdto point at the new folder structure
Going forward: every experiment, config change, or methodology lesson gets a timestamped entry (## HH:MM — title) in the current day's file. New day = new file. Current State in README.md gets updated in place whenever production state changes.