Quant panel - suggestions report (2026-08-01)
Five outside experts (volatility/VRP, short-side microstructure, chief risk, validation methodology, new-alpha/data strategy) filed 25 suggestions; three adversarial reviewers killed 2, forced 8 near-duplicates into 3 merged items, and attached binding revisions - folded directly into the items below. Plain language throughout: every term of art gets a one-clause translation at first use.
1. Executive summary
Nineteen items survive. The top 5, ranked - what, why now, cost, kill criterion:
- Fix the go-live leverage diff before December (C1) - the pre-specified live config caps the wrong knob (the vol-target multiplier, not actual market exposure) and would deploy a never-backtested ~0.14x-exposure, ~4%-volatility book while saving zero margin interest; 1-2 days plus twin backtests; adopt only if the corrected twin matches the validated baseline on >90% of days with no overfitting-score or per-window regression, else keep the diff as written and pre-write the honest under-vol disclosure.
- Persist the option IV term structure the snapshotter already holds (B1) - it is fetched daily and thrown away at zero extra API cost, and every unpersisted day is a permanent hole; ~1 day; retire the stream if after 60 trading days it recovers under 5 expiries median per name.
- Fix the VRP pessimistic twin's quote basis (A1) - the sleeve's honest-cost ledger printed an impossible negative spread cost because closes are priced off a different quote snapshot than the decision; ~0.5 day; if impossible values persist post-fix, the cost series carries an uncertainty band and every downstream constant comes from its conservative edge.
- Null-calibrate the kill instrument (C2) - run the shop's own overfitting harness on books with no skill (SPY, ~20 seeded random 50-name books, timing-averaged composites) so "PBO 50%" finally has a printed scale; 1-2 days setup plus a weekend of compute; if the no-skill distribution centers near 50% with correlation near 0, the current absolute thresholds are vindicated and the tranching door closes permanently.
- Make the multiple-testing ledger append-only (C3) - the trial count behind the deflated Sharpe is currently deletable and incomplete (it drifted 0.665 to 0.613 with no strategy change); 1-2 days; if the completed count plus era-scoping move the number by <0.05, keep only the append-only rule and record the rest inert.
Next tier, all deadline-bound before 2026-12-01: the merged fleet risk amendment (C7 - short-vol must never be sized by its calm-window volatility), the December governance package (C6 - scale-up ladder with actual numbers, evidence-or-exit sunset), and the sequential retire-monitor registration (C8). Cheapest strategic buy on the list: pin the skew activation test now (B4), ten months before the data can speak.
2. How to read this
Everything here is a pre-registerable experiment or an instrument fix, not a trade. The backtest's only trustworthy job is falsification (killing configurations), while the forward paper track is the only judge of live expectation. Nothing below touches a live pinned rule mid-gate: amendments land only where a rule has never consumed data or where implementation demonstrably contradicts its own registered intent.
Standing caveats on every number: (1) every absolute Sharpe (return per unit of risk) carries the timing-luck spread - the frozen-window 0.795 is the rank-1-of-8 start-day draw; across start days the same signal gives mean 0.09, range -0.32 to 0.80. (2) Paper fills print at NBBO (the market's best quoted price), so the paper track cannot measure execution cost - hence the pessimistic twin. (3) The December gate is an operational-competence gate authorizing a capped 25% experiment, not alpha proof; NOT-PASS is the modal outcome and weak evidence.
3. The suggestions
Track A - Do now, cheap
A1. Same-snapshot fix for the VRP pessimistic twin (derivatives PM; 3/3 survive). The twin re-prices every paper fill at the worst side of the quoted spread. Its close-side cost uses a re-fetched quote from a different moment than the decision, mixing spread cost with market drift: the 07-31 close printed a pessimistic price BELOW the decision mid - impossible from one quote - and the ~0.03/share noise equals the ~0.075/share measurand VRP economics hang on. Fix in execution/vrp_daily.py: price the close from the same quote read as the decision (the open path already does). Test: unit-assert worst-side >= mid from one snapshot; ~5 cohorts forward, zero negative drags expected. Kill: persistent negative drag = unstable feed - the series gets an uncertainty band and downstream constants use its conservative edge; drag >=50% of credit even at wider structures = pre-registered family-level kill of defined-risk spreads at retail friction. ~0.5 day. Red team: forward-only; annotate the 4 existing rows basis-contaminated.
A2. Deployment-decomposition telemetry + shadow gross (CRO; 3/3). Exposure is set by three stacked layers - Kelly (edge-conditional deployment fraction), the vol target (risk normalizer), caps - and sits on a knife edge (3x Kelly 0.419 vs vol-target 0.418). One report-only daily line: which constraint binds, plus the exposure the currently-registered go-live config would produce - making December's deployment-vs-alpha split mechanical. All inputs owned; lockstep test pins it to the construct.py formula. Kill (vacuity): 63 trading days of unchanged binding label with shadow within 10% of actual collapses it to one digest line. ~1 day. Red team: the shadow must track whichever go-live diff is currently registered - never a dead counterfactual.
A3. Short-flow to short-interest nowcast - one-shot falsification (short-side specialist; 3/3). The twice-monthly short-interest print (total shares sold short) arrives ~3 weeks stale; IF daily FINRA short volume carried positioning information, cumulative flow surprises would predict the next print's change. The specialist's measurements say it mostly will not (22-54% venue coverage; a 57% median short-marked share is market makers shorting to fill customer BUYS - liquidity provision, not bearish bets), so one cheap shot settles it forever. Data: owned - 2,002 daily files + 205 partitions; no returns anywhere, zero multiple-testing cost. Test: Spearman rank correlation (do predicted and realized changes rank names similarly) per window; parameters fit 2018-21, verdict on untouched 2022-26. Kill: median out-of-sample correlation <0.20 - record the null, demote the stream permanently, bar future proposals from citing daily short volume as positioning without venue-complete data; >=0.20 earns only "eligible as timing refinement" behind its own registration. 1-2 days.
A4. Publication-lagged days-to-cover exclusion battery (short-side specialist; 3/3). When shorting is expensive or crowded, pessimists are locked out and the price reflects optimists only; a momentum winner with high days-to-cover (short interest / daily volume - days shorts need to buy back) is often a squeeze the long book buys late. Tail-avoidance, not new alpha - the exact consumer the short-interest registration already named. Data: owned 205 frozen partitions (2017-12 on) + golden store + point-in-time membership. Mandatory first: the engine hook has a verified ~7-business-day lookahead (settlement date used as knowledge date, engine.py ~L780) - add a publication-lag knob. Screen pinned at DTC>10 (inherited, not tuned); one sensitivity (p90 7.9) reported, never selected. Test: hash-frozen memo; vacuity count BEFORE any return is seen (<1 name-change/quarter = vacuous, capture-only); only then walk-forward A/B + both CPCV modes under the full SHIP gate. Kill: overfitting score worsens or any window regresses >=0.05 Sharpe. Modal outcome: vacuous-or-inconclusive - a legitimate cheap closure. 2-3 days. Red team: pin exact baselines and bootstrap details; earliest ship post-December (a selection change now resets the clean clock).
A5. TRACE credit-deterioration probe (new-alpha strategist; 3/3). Bonds of leveraged firms often reprice credit stress before the stock breaks; the only long-only use is exclusion (the GE-2018 shape). Honest prior: probably vacuous. ~1-day probe: does FINRA's auth-free Query API (already used in-house) expose per-issuer end-of-day bond prices from TRACE (the bond trade tape) with >=3 years depth? No scraping, no paid feed - pre-declined. Kill ladder: probe fails - drop and record; passes - bond-to-ticker mapping (1-2 days), vacuity check before any return (>=1 name-change/quarter), then hash-frozen battery + SHIP gate. Red team: pin "leveraged issuer" BEFORE probing; the vacuity report must show overlap with names the emergency-exit machinery would have exited anyway; BACKFILLABLE, acquire at need, no standing cron.
A-salvage. Adjusted-opens audit (rescued from the killed overnight/intraday item; all three reviewers endorsed it standalone): verify overnight x intraday returns reproduce close-to-close per name-day; census |overnight| > 20% days against corporate actions; quarantine names failing >0.1% of days. Vendor-adjusted opens are the least-trusted field in any feed; TCA arrival math reads them. ~1-1.5 days of ops hygiene, no registration.
Track B - Capture now, test later
B1. Persist the per-expiry IV summary (term structure) (derivatives PM; 3/3). Implied volatility (IV - the market's priced expectation of future movement) across expiries is the standard measure of event-vol richness and vol carry, predicting OPTION-strategy returns - explicitly not the killed stock-alpha variant (that kill's own text says the effect predicts option/index, not stock, returns). The snapshotter already fetches every expiry and persists one (data/options_skew.py ~190-212): zero extra API calls, each unpersisted day irreversibly lost. Persist a compact per-expiry row (date, name, expiry, days-to-expiry, at-the-money IV, 25-delta put/call IV, counts) - a few hundred KB/day, ~100x below the rejected full-surface option. Also: verify whether the feed exposes per-contract daily volume (the only crowding proxy given no open interest); add underlying spot. Red-team conditions: register in CAPTURES.md BEFORE accruing, consumer class pinned as option-strategy research only; pin the first-consumer test at capture time (slope definition, date, threshold) with the honest power note and post-close quote caveat. Kill: <5 expiries median per name at 60 trading days - recorded vacuous, retired. ~1 day. Highest value-per-effort on the list: delay is the only expensive choice.
B2. Option-implied borrow-fee wedge (merged: short-side specialist + new-alpha strategist; one registration required). Put-call parity (the no-arbitrage relation tying puts, calls, stock, cash) bends when a stock is hard to borrow: market makers hedging puts must short it and pay the fee, so calls cheapen versus puts by roughly the fee. Deriving this "wedge" from the already-paid-for chains capture converts the binary borrow flag - degenerate on the flagship universe (0 of 889 names hard-to-borrow) - into the graded fee measure the shortability registration says is missing, and cleans the skew series (which confounds crash fear with borrow cost). This inverts a recorded kill legitimately: those measures were killed AS alpha signals; here the artifact IS the measurement, no returns touched. Method: conversion-reversal BOUNDS for dividends; ranks and innovations vs each name's trailing ~20-day wedge (levels are ~a dividend detector). Validation, red-team rebuilt: the pinned hard-vs-easy separation test CANNOT RUN (empty positive class in-universe); replace with feasible ground truths, each with stated n - margin-requirement tail (n=4, reported not binding), debounced flip-event elevation pooled over accrual, days-to-cover-tail concordance, rank-stability floor. One memo, thresholds pinned before any computation. Kill: fail the suite - instrument dead, raw chains keep accruing; no ground truth reaches its pinned n in a dated window - UNVALIDATED-USABLE-AS-RANKS-ONLY. Consumption: enters A4's battery memo only once actually frozen; signal use behind the full SHIP gate. 3-4 days; may retire the "IBKR account for fee levels" path for free.
B3. SEC fails-to-deliver - acquire and freeze (short-side specialist; 2 REVISE / 1 SURVIVE). Fails-to-deliver (shares sold but not delivered at settlement) mark names where the borrow plumbing is visibly breaking - the same overpricing tail as days-to-cover, measured from settlement stress rather than positioning, with deep free history (~2004+). Verified absent from the reject list. Free immutable twice-monthly SEC files; probe first (cadence, lag, and the historical aggregate-size floor that may thin Russell-1000 coverage - record per-year fill rates first); drop at zero cost on paywall/bot-hostility; knowledge date = publication date; BACKFILLABLE, no cron. Red-team revisions, binding: pin the screen from an EXTERNAL convention (e.g. Reg SHO threshold criteria) or freeze the memo BEFORE the bulk history lands - never a percentile chosen after the distribution is on disk; inside A4's battery it is REPORTED-ONLY and can never substitute for the primary (an FTD pass with a DTC fail ships NOTHING); counted as a second trial (N+1) at freeze. Kill: >=90% overlap with DTC-based selection changes - redundant, archive-only. 1-2 days.
B4. One pinned skew/VRP activation test (merged: derivatives PM + new-alpha strategist; single binding registration required). The skew stream's activation criterion - "escalate if a residual long-only signal appears in 12-18 months" - names no test, so the mid-2027 evaluator could pick the definition that flatters the preferred outcome. Close it now, while the track is too short to peek at (the placebo-pinned-at-day-22 precedent). Under test: steep put skew (downside protection priced rich - a fingerprint of informed sellers) predicting poor stock returns, versus the rebuttal that it is mostly a borrow-fee artifact a long-only book cannot monetize anyway; the shop owns data on both sides. Pins: ONE binding primary - the Russell-1000 weekly cross-sectional rank-IC (information coefficient: rank correlation between today's signal and next week's returns), chosen for power; the momentum-top-100 de-selection spread (the only sanctioned consumer) as pinned secondary; graded constraint conditioning (days-to-cover deciles, margin tail, B2's wedge if validated) instead of the binary flag - degenerate here, and a branch that can never fire would launder "artifact not confirmed" out of a test that could not confirm it - with an UNDECIDABLE-ON-THIS-UNIVERSE clause; one primary variable or a multiple-comparisons bar (|t|>=2.4 for any-of-3); power printed (screen-grade, not SHIP-grade); dates pinned (2027-06, repeat month-18, nothing between). Thresholds gate SPENDING (buying retail chain history - legitimate since the scarcity retirement), not truth. Kill: below threshold at both checkpoints - stream demotes to insurance, no relitigation. ~1 day; lands as a dated CAPTURES.md amendment.
B5. Turn-of-month forward attribution (new-alpha strategist; 2 REVISE / 1 SURVIVE). The headline Sharpe's rank-1/8 start-day luck is either turn-of-month structure (returns clustering at month boundaries - plausibly pension/payroll flows) or plain schedule luck - and the live day-1 book is ALREADY running the decisive forward experiment, unlabeled. Pin the decomposition now; report-only, never gates. Red-team revisions: decisional window = 12 clean months FROM THE PIN DATE - the ~36 accrued days appear only as a flagged prelude, never in the judged statistic; phase pinned as last-1 + first-3 trading days, lockstep-tested; bootstrap block length and confidence level pinned; attestation that no phase-conditioned read preceded the pin; power printed (~48 turn days/year gives t ~1 at month 12 - do not overclaim). Three pre-pinned outcomes: CONFIRMED (interval excludes zero, right sign); REFUTED (signed-negative point estimate - schedule-luck reading adopted); INSUFFICIENT (positive, interval spans zero - clock extends to a pinned month-24 read, no parameter changes). Deployment planning may adopt the conservative across-schedule mean (~0.09) at ANY time as policy, decoupled from the verdict. 0.5-1 day.
Track C - Methodology upgrades
C1. Go-live leverage diff: cap gross exposure, not the scalar (CRO; 3/3). The pre-specified December change (max_leverage 3.0 to 1.0) caps the vol-target SCALAR - the multiplier levering the Kelly-sized book to its 12% vol target - not gross exposure (market value held per dollar of equity). Verified: with Kelly ~0.14, the literal diff deploys a ~0.14x-exposure, ~4%-vol book no backtest validated, saving zero margin (the paper book borrows nothing at 0.42x; the diff's own written rationale is exposure-denominated). Amendment: "no borrowed dollar" - scalar cap = min(3.0, 1.0/Kelly-gross): the validated composite everywhere the account lives, clipped only above 1.0x exposure. Not gate-shopping: visible from the formula alone, precedes the gate and any real dollar, and INCREASES risk in losing streaks - the ratified implementation-matches-intent class. Test: three frozen-window twins (baseline / gross-cap / literal), adoption criteria committed BEFORE the first twin runs. Kill: adopt only if the gross-cap twin differs on <10% of days, survivorship-free PBO <= baseline, no window regresses 0.05 Sharpe; else keep the diff and pre-write the honest disclosure. Red team: a NEW engine cap mode - the re-baseline rule applies; runs logged as calibration. 1-2 days; before 2026-12-01.
C2. Null-calibrate the falsification engine (methodologist; 3/3). The central kill statistic - PBO (probability of backtest overfitting: how often the best-looking in-sample path lands below median out-of-sample) - is a nonstandard single-config variant moving in 12.5-point quanta, and negative in-sample/out-of-sample correlation is partly arithmetic for episodic strategies (hot episodes in test blocks are, by construction, missing from the train metric). Nobody has measured what the harness prints for NO skill - every kill threshold is an absolute number on an uncalibrated scale. Run the pinned harness on (a) SPY, (b) ~20 seeded random 50-name books, (c) 2-3 timing-averaged composites - probe (c) decides whether the tranching kill was partly a statistic artifact. Reading pre-registered BEFORE any run: calibration only; all existing verdicts stand; future thresholds quoted as null percentiles. Kill (of the artifact hypothesis): null centers near 50% within one path-quantum, correlation near 0 - thresholds vindicated, tranching door closes permanently; only if (c) shows systematic punishment of timing-averaged books does a tranching v2 earn a NEW hash-frozen registration - eligibility to register, never to ship. Red team: null books need small test-tier strategy classes (production untouched); ~25 runs is plausibly a weekend of home compute; pin the refutation statistic exactly and carry the 20-draw resolution limit. 1-2 days setup; land before any sleeve-#4 battery.
C3. Append-only trial registry for the deflated Sharpe (methodologist; 3/3). The deflated Sharpe (observed Sharpe judged against the best of N lucky tries) is only as honest as N - which today comes from a deletable directory (evaluate --delete erases tracelessly), misses ~138 CPCV-only runs and trials living in config comments, pools dispersion across engine eras (a dead engine's 1.66 Sharpe inflates it), and silently drifts (0.665 on 07-21, 0.613 on 08-01, same strategy). Build: one committed JSONL registry (date, sleeve, family, era, instrument, headline, verdict); mechanical backfill; deletions become tombstones; lockstep test asserts registry contains the evaluations directory. Report as a sensitivity band (axis-set vs raw N); era-scoped dispersion ONLY inside the band next to the pooled number - era-scoping alone is a flattering change smuggled inside an integrity fix. Mechanical family labeling; raw row count always reported; every future battery cites the row-count at freeze; per-trial daily returns accrue forward. Kill: refinements move the number <0.05 - record inert, keep only the append-only/tombstone rule (the actual fix, free to maintain). 1-2 days.
C4. Bound the stateful-purge leak (methodologist; 2 SURVIVE / 1 REVISE). The purge (buffer of days deleted around test blocks so training cannot touch them) covers the 252-day feature lookback - but the in-sample backtest runs contiguously THROUGH interior test blocks carrying slow state: the ~504-day Kelly ledger and the kill-switch high-water mark. Test-block returns therefore sit inside in-sample sizing - a flattering-direction leak on the primary kill metric, never measured. Red-team redesign: run the 2x2 - Kelly window {24mo, 12mo} x purge {252, 504} - since neither variant alone isolates the leak; only the INTERACTION exceeding a pre-pinned materiality bar (one path-quantum of PBO, 0.15 of correlation, restated in realized path counts) extends the purge floor to cover stateful estimators (one-line change + re-baseline). Main effect alone = design sensitivity, not leak; all cells inside the bar = BOUNDED-IMMATERIAL, recorded. The purge-504 arm is the cleaner probe and ranks primary. Production Kelly untouched. ~1 day + overnight compute.
C5. Beta-hedged companion for the placebo ensemble (methodologist; 1 REVISE / 2 SURVIVE). The pinned placebo null - 200 random books answering "is the live Sharpe skill or luck?" - carries full market exposure (beta ~1) while the live book runs ~0.49: in an up-tape the percentile handicaps the incumbent for reasons unrelated to stock picking, flatters it in a down-tape - exactly as it becomes decision-grade (~60 days, late September). Add a REPORT-ONLY companion: subtract each book's beta x SPY return from both sides, print both percentiles; the raw pinned one stays the only binding number - changing the pinned null mid-window would be the actual sin. Red-team revisions: register NOW, before decision-grade; the dated append records what was observed at registration (window, tape direction, current percentile, and that the companion TODAY flatters the incumbent) so pre-dating is auditable; divergence reading neutralized - it means only "the beta confound is material; neither percentile is privileged"; each render carries the beta-estimation caveat (standard error ~0.15-0.25 on short windows). Kill: <15-point divergence over 6 months - inert, retired. <1 day.
C6. December amendments: scale-up ladder, sunset, economics companion (CRO; 2 REVISE / 1 SURVIVE). Three governance holes at the capital layer, closable now precisely because the window is low-power and December's Sharpe outcome is effectively known - nothing can be tuned to flatter. (a) Ladder: "review at +63 trading days" names NO criteria; pre-register rungs - zero endogenous halts, real-fill transaction costs median <=50bps, live-vs-paper tracking error inside a band, no DEGRADED; a failed rung holds or steps down MECHANICALLY (pin which fraction). (b) Sunset: if by 18 months of clean clock the gate has never passed AND placebo <60th AND active Sharpe <=0, retire - closing the zombie zone between retire (-2 sustained) and pass (~+2.3). (c) Companion: growth per unit of exposure deployed. Red-team revisions, binding: every named number gets an actual value IN the append (an unpinned band is not a pin); publish the sunset's power table (its active-Sharpe leg fires on a true-0.3 strategy ~36% of the time - accepted in print); (c)'s privileged "deployment fact" reading STRUCK (drafted knowing the live exposure and modal outcome) - symmetric interpretation plus disclosure of what was observed at drafting; name the assumption that the paper lane keeps trading as comparator, and what happens if it dies. Retirement via two DEGRADED first makes (a)/(c) moot - pre-agreed, never a reason to loosen DEGRADED. 1-2 days; before 2026-12-01.
C7. One fleet risk amendment: stress-loss denominators, correlation-evidence gate, pinned scenarios (merged: derivatives PM + CRO, three overlapping drafts; reviewers unanimous on ONE artifact). All three sleeves are the same crash bet - long equities (beta ~0.49), the long-equity control, short SPY puts. The pinned allocation rule (equal risk = inverse trailing 126-day realized volatility of each sleeve's equity curve) would systematically OVER-allocate to short-vol: an insurance seller's calm volatility (~1%) hides its tail - the classic short-vol Sharpe illusion - and calm-regime correlation is the wrong number for crash coherence. Fixable without gate-shopping only now, before any cross-sleeve dollar exists. One dated keystone amendment to fleet.yaml: ONE denominator convention - each sleeve's own structural worst case (VRP: open quantity x width x 100 / equity, what its safety gate already computes; weight sleeves: exposure x ONE pinned crisis-vol convention, the draft's either/or resolved at registration); a correlation-evidence gate - no diversification credit (allocate as if correlation = 1) until >=126 overlapping live days AND one joint stress observation (>=5% SPY drawdown), with daily-P&L correlation pre-committed as NON-evidence at VRP's quantization scale; a pinned scenario table (2020-03 and 2008-10 worst 5 days, VIX-doubling) with a lockstep test so scenarios cannot drift. Justification from payoff structure ONLY - no live-curve numbers. Render ONE digest line; the standalone weekly crash report was struck as born-vacuous at today's sizes (VRP max loss ~$200 vs ~$34k fleet equity) and is built only on a pinned trigger. Kill: after two joint stress windows, identical rankings and allocations within 10 points - redundant, revert and record. 1-2 days; before any cross-sleeve allocation decision.
C8. Sequential retire monitor for the real-money era (methodologist; 3/3). Fixed-horizon tests are the wrong tool for a slow judge: the December Sharpe criterion has ~12% power, and the retire rule needs live Sharpe ~-1 sustained - real money could ride a dead edge for years inside the -20% kill-switch. An e-process (a sequential evidence measure you may check daily and stop on at any time, without the repeated-peeking penalty that invalidates ordinary p-values) on daily ACTIVE returns (live minus SPY, dodging the beta confound), pre-registered NOW for an era with zero accrued data. Hash-freeze before 2026-12-01: statistic family and prior pinned BEFORE any simulation (else the calibration sweep becomes a design search that passes its own gate); publish the FULL simulated operating-characteristics table (time-to-signal at true Sharpe -0.5/0/+0.3/+0.65, false-halt rate). Adoption gate: median time-to-halt at true zero above ~2.5 years, or false-halt at a real 0.65 edge within 2 years above ~15% - record the negative, do NOT adopt. Red-team warning in print: the gate probably fails on the arithmetic (~14 years median at a conventional threshold, per one reviewer) - budget as a likely pre-registered negative that permanently answers "why no sequential test?", at maybe 2x the claimed 2-3 days. Worth buying is Dan's call (section 6).
Track D - Post-gate candidates
D1. VRP v2 spec: restore the vol content, amortize the friction (derivatives PM; 2 REVISE / 1 SURVIVE). The variance risk premium (the persistent gap between what hedgers pay for index puts and what they prove worth) is harvested only with actual vega (sensitivity to volatility) and holds long enough to amortize friction. v1's $1-wide spread has ~no vega, burns ~40% of credit per 8-9-day cycle, and its 21-days-to-expiry exit guards a blow-up zone a capped-loss spread does not have. v2: entry ~45 days out, short strike -25 to -30 delta, width 2-3 registered JOINTLY with the risk budget so the contract count sits mid-interval (v1 sits exactly at 2.00 budget multiples - an $8 equity dip halved the position), exit at 50% of max profit or a 7-10-day time stop, entry filter credit >= k x measured drag. Red-team revisions change the shape: hash-freeze the spec NOW, not "draft now, freeze after v1" (an open draft would absorb v1's outcome); exactly ONE deferred slot - the drag constant, entered mechanically at activation via a formula pinned today (median round-trip drag on the POST-A1-fix twin over v1's full window; minimum-observation rule; conservative edge if the band is wide); jointly re-register the safety gate's per-structure loss cap - as drafted, width 2-3 breaches the untouched 2% cap and the sleeve sits structurally flat (width 2 x qty 1 = exactly 2.0% at $10k; any dip rejects every open); dated appends only, none once v1's result is visible; activation only at v1 resolution (~2027-01) plus the fleet cap. Gate: same 126-day scaffold, twin from day one, a drawdown line that can actually bind (~5x worst single-cohort loss). Kills: pessimistic-curve net credit <=0; time-stop closes >80% of cohorts; drawdown breach regardless of Sharpe. Pre-committed ceiling in the spec's own words: 126 trading days cannot prove the premium exists - only that the harvest is not cost-dominated and operations are clean (the Sharpe leg is a ~5-12%-power coin flip; the structural falsifiers are the real gate). 1-2 days of writing.
D2. The tranching door (conditional, not a proposal). Only if C2's probe (c) shows the overfitting statistic systematically punishes timing-averaged books does a tranching v2 earn the right to a NEW hash-frozen registration - eligibility to register, never to ship. Otherwise the 2026-06-10 kill stands permanently.
4. Sharpest findings about the current system
Deduplicated; file references where given.
- The go-live leverage diff is a category error (
config/settings.yamlgo_live;portfolio/construct.py): it caps the vol-target scalar, not exposure - verified live (gross 0.419 = Kelly 0.140 x scalar 2.996). As written, December deploys a ~0.14x, ~4%-vol book no backtest validated. - December's Sharpe criterion is effectively decided: live Sharpe -2.34 at n=34 is statistically nothing (standard error ~2.7), but passing needs ~+4.0 annualized over the remaining days - December's live content is the operational sections, as the memo pre-committed. The negative control (+4.5% since 07-13) beating the incumbent is the noise-floor demo working.
- Governance may bind before December: the DEGRADED retire rule is a marginal call by the September/October runs; two consecutive fires retire the incumbent (
results/diagnostics/oos_verdict_history.jsonl). - VRP width-1 economics are cost-dominated by its own measurements (
data/state/vrp/pessimistic_ledger.jsonl): ~0.07-0.08/share round-trip friction vs 0.18 credit, near-zero vega, profit target near-unreachable in both live cohorts - the counterparty collecting the spread is the market maker. - The twin's ruler is broken (
execution/vrp_daily.py): the 07-31 close printed a negative spread cost, impossible from one quote. - Two sizing knife edges: VRP sits exactly at 2.00 budget multiples at $10k (an $8 dip halves the book); momentum sits at 3xKelly 0.419 vs vol-target 0.418.
- The fleet is one crash bet (
config/fleet.yaml): all three sleeves short the same left tail; the pinned inverse-realized-vol allocator would overweight short-vol for exactly the wrong reason. - PBO has never been calibrated (
backtest/cpcv.py:_compute_pbo): nonstandard, 12.5-point quanta - the eligible/killed boundary is ONE path flipping - and negative in-sample/out-of-sample correlation is partly arithmetic for episodic strategies. - Multiple-testing debt sits at the hurdle: 274 eval files, 113 axis-sets; luck hurdle 0.455-0.504 vs honest observed 0.524; deflated Sharpe 0.613/0.528 by counting choice; the ledger is deletable, incomplete (138 CPCV files uncounted), era-pooled.
- The Kelly window leaks through the purge (
settings.yaml;cpcv.py): ~504-day ledger vs 252-day purge - flattering direction, magnitude unmeasured. - The placebo null is beta-mismatched (
backtest/placebo.py): beta ~1 books vs live 0.49 - regime-confounded exactly as it becomes decision-grade (~late September). - The binary borrow flag is degenerate on the flagship universe (
data/shortability/): 0/889 hard-to-borrow at rest; the graded variation lives in margin requirements (GME 100/MRNA 60/CAR+TSLA 50) and days-to-cover (median 4.3, p90 7.9). Pre-register a 2-day flip debounce NOW - two of the first four flips look like vendor noise. - Daily short volume is not positioning: 22-54% of tape volume; a 57% median short-marked share is market-maker liquidity provision - often BUYING pressure.
- A real lookahead in the days-to-cover hook (
backtest/engine.py~772-790): settlement date used as knowledge date vs ~T+7-business-day publication. - Already-paid-for data holds two upgrades: the borrow-fee wedge inside the chains capture, and the discarded IV term structure (
data/options_skew.py~190-212). The skew activation criterion names no test (CAPTURES.md) - a June-2027 result-shopping fuse - while the live day-1 schedule runs the decisive turn-of-month experiment unlabeled. - Kill-list entries are being read wider than their records: the index-reconstitution kill covers S&P ADDITIONS, not small-cap deletions; the micro-cap kill covers the ranking signal, not event studies.
- Ops residue for real money: single broker across all sleeves; one home machine for backup and the monthly integrity tripwire; no key-rotation procedure in the RUNBOOK. A rotation drill belongs on the pre-December checklist.
- Credit where due (three experts independently): the pre-registration discipline is stronger than most institutional shops'. The apparatus is currently stronger than the strategy - the right way around.
5. Rejected by the panel
The reject ledger is a first-class output.
- Overnight-vs-intraday momentum co-signal - KILLED 3/3: a transform of the same owned price data feeding another cross-sectional ranking inside the thrice-ratified exhausted large-cap space; a flow story attached to OHLCV does not make the data new; buying an expected null costs a trial-ledger entry plus an engine-touch re-baseline. The proposer's own prior was KILL. Salvage: the adjusted-opens audit ships standalone (A-salvage).
- Russell-2000 deletion-reversal battery - killed by ONE reviewer (statistical), REVISE from two; per panel rule a single-reviewer kill lands here by default, and I am explicitly NOT exercising the rescue option: the power argument is unrepaired by the other reviewers' fixes. Six annual event cohorts each share one market draw (effective sample ~6); clearing the pre-registered confidence bound needs a ~7-10% net effect against literature-scale low single digits (power ~10-25%); the unpinned bootstrap unit (name vs cohort) flips the verdict by itself; the sign backstop is 34% false-positive at true zero. A foredoomed one-shot would be enshrined as fake-definitive by the never-relitigate covenant - worse than not running. Revival condition recorded: demonstrate >=50% power under cohort-level inference FIRST, plus the other reviewers' fixes (coverage floors, pessimistic treatment of deleted names with no owned prices, the knowledge-date fix, adjudication of the S&P-additions kill-list entry). The IWM membership asset stays owned and idle; killing the battery spends nothing.
Struck-in-part (folded into survivors): the standalone weekly fleet crash report (born vacuous at today's sizes; scenarios survive in C7) - the "deployment fact" privileged reading (pre-drafted excuse; symmetric reading required, C6) - the placebo companion's original divergence reading (same defect; neutralized, C5) - the hard-vs-easy borrow validation for the wedge (empty positive class in-universe; replaced, B2) - the binary borrow-flag conditioning in the skew test (a kill branch that could never fire; replaced, B4) - VRP v2's "draft now, freeze later" sequencing (open-draft result absorption; replaced by freeze-now-with-one-slot, D1).
6. Decisions only Dan can make
- Budget and order. ~25-30 person-days total. The ranking is in section 1; the split between the pre-December governance cluster (C1, C6, C7, C8) and the research-now cluster (tracks A/B) is the operator's call.
- Risk-appetite numbers. The fleet crash budget (what fraction of fleet equity may die in a 2008/2020 week), the ladder's tracking-error band, and accepting the sunset's ~36% retirement probability at a true-but-modest edge are appetite settings, not analysis.
- The sequential monitor's price. C8 is a likely pre-registered NEGATIVE costing maybe 4-6 days; buying a permanent honest answer to "why no sequential test?" is a taste decision.
- The 2027 spend gate. B4's thresholds gate the decision to buy retail options history; the dollars are yours.
- Whether VRP v2 runs at all. If the post-fix twin says friction stays >=50% of credit even at width 2-3, the family is untradable at retail friction and D1 dies without activating.
- Compute and attention. C2's null batch is a weekend of the home machine; every standing report line is a permanent claim on the shop's scarcest resource - one operator's attention.
- The December decision itself. Every item here makes it more legible; none of them makes it.
Appendix: the five experts
Derivatives PM (volatility). Found VRP's economics inverted by its own measurements - width 1 stripped the vega and left ~40% of credit as friction - while calling the pre-registration discipline professional-grade ("the apparatus is stronger than the strategy, the right way around"). Contributed the twin fix, the term-structure capture, the v2 spec, and the one-crash-bet fleet warning.
Short-side microstructure specialist. Measured rather than opined: the binary borrow flag is degenerate on the flagship universe (0/889), daily short volume is market-maker inventory, and the salvage is graded and mostly owned - days-to-cover, the margin tail, the parity wedge. Caught the engine-hook lookahead. Every proposal is exclusion or measurement, never a harvested premium; honest expectation vacuous-or-inconclusive, run cheaply to closure.
Chief risk officer. Called this the most honestly-governed solo book they had reviewed, then went after the seams pre-registration has not reached: the leverage category error (the headline), the criteria-free scale-up review, the retire/pass zombie zone, the inverse-vol fleet trap, the ops residue. Noted December's Sharpe verdict is effectively known - making now the cheapest moment to pin everything around it.
Validation methodologist. All calibration, no alpha: the kill statistic has never been measured against a null, the trial ledger is erasable and undercounts, the Kelly window leaks through the purge, the placebo null is beta-confounded, the retire monitor is slow by years. On timing luck: report-only is correct; the ensemble idea is already on the kill list, and the only legitimate reopening path runs through null calibration.
New-alpha strategist. Scoped the exhaustion doctrine precisely (cross-sectional price+fundamental ranking on large caps) and hunted inside owned data: the wedge in the chains capture, the 2027 activation holes, the unlabeled turn-of-month experiment. Two of six proposals died - one by their own stated prior, one on power - roughly the hit rate an honest idea pipeline should expect.