Micro-cap momentum battery — PRE-REGISTRATION v1 (2026-07-22)
Binding document. Everything above the sentinel line is hash-frozen by
tests/test_docs_consistency.py the moment this merges. Amendments are
allowed ONLY before the first battery run, as a dated superseding section
appended below the sentinel plus an explicit hash update in the test — a
recorded act, never a silent edit. After the first run, the results stand
against exactly this text. Choosing any value below after seeing micro-cap
backtest results would be the gate-shopping this document exists to prevent;
as of this writing, zero micro-cap backtests have ever been run on this
apparatus.
Evidence base: research/2026-07-21_microcap_data_feasibility.md
(resolution + build log, workstreams 1–3) and
data/microcap/cost_table.json (measurement artifact).
1. Hypothesis (exact, no signal tuning)
The flagship's incumbent registered signal — 252-day momentum, skip 5, blended 50/50 with 21-day smoothness, features shifted one day — applied to a small/micro-cap universe, survives honest (upper-bound) costs because post-publication factor decay concentrates where arbitrage capital cannot deploy (McLean–Pontiff). No signal parameter may differ from the flagship's production values. A variant signal is a different hypothesis requiring its own pre-registration.
2. Universe (pinned)
- Membership: point-in-time Russell 2000 via the Wayback-IWM record
(
data/constituents/r2000_membership_iwm.parquet); a name is a member on date d iff present in the most recent snapshot ≤ d (monthly cadence; the 9 unrecoverable snapshots are tolerated holes). - Identity: (symbol, membership era) — never symbol alone (BBBY ticker-reuse finding, build log).
- Eligibility at selection (non-performance justifications only): price ≥ $1.00 (exchange delisting-standard threshold), ≥ 60 prior daily bars in the owned stores (estimator/covariance floor), and the borrow screen of §5.
- Selection: top 50 eligible members by the signal (flagship's N), monthly on the first trading day, daily risk management between — the flagship's registered cadence, unchanged.
- Construction: identical to the flagship's registered production construction (HRP weighting, half-Kelly, 10% position cap, 35% sector cap, 12% vol target VIX-scaled, 5% daily turnover cap frequency-scaled, 20% kill-switch). Universe, costs, delisting handling and the borrow screen are the ONLY differences from the flagship.
3. Window (pinned)
2020-01-01 → 2025-06-30. The survivorship-clean span (SimFin's delisted coverage begins ~2019-20; feasibility memo OQ#3). Pre-2020 runs are forbidden under this registration — extending backward requires a superseding version with a measured Wayback micro pilot or an attached survivorship caveat on every number.
4. Costs (pinned upper bounds — false kills over false ships)
Per-name flat round-trip assigned by trailing-252-bar median dollar ADV at
selection, from the measured p75 upper-bound table
(cost_table.json, Abdi-Ranaldo primary, resolution floors documented):
| $ADV bucket | round-trip bps (binding) |
|---|---|
| < $250k | 414.2 |
| $250k–1M | 194.2 |
| $1M–5M | 120.0 |
| $5M–25M | 103.7 |
| ≥ $25M | 114.4 |
(= p75 AR spread + Almgren-Chriss impact at 1% ADV participation, η = 0.1 pinned — uncalibratable without fill data.) Sensitivity at ±50% of this table is REQUIRED in the battery output and is informational — except: a configuration that passes only under −50% costs is a KILL.
5. Borrow/crowding screen (pinned)
Exclude at selection any name whose most recent FINRA consolidated
short-interest settlement (≤ selection date, data/short_interest*) shows
days-to-cover > 10.0 — the Muravyev–Pearson–Pollet borrow-fee-artifact
class is screened OUT, never harvested. Names with no SI record are
INCLUDED (missing ≠ shorted).
6. Delisting returns (pinned)
If a HELD name's price series terminates while held (no further bars), the final holding return takes an additional −55% terminal haircut (the harsher Shumway-class Nasdaq figure, applied uniformly because our data cannot distinguish deaths from acquisitions; over-punitive for acquisitions — stated, and in the false-kill direction). A no-haircut run is REQUIRED in the output as sensitivity, and is non-binding.
7. The battery (pinned)
On the stores as they exist at run time (microcap store + data/raw; the optional 306-name Tiingo alive-gap fetch may be added BEFORE the run — adding data is not tuning):
- Survivorship-free CPCV: n_groups = 6, k_test = 2, embargo = 5,
purge = max(252,
min_required_purge(config)). - Walk-forward with bootstrap CIs.
- Start-day sweep — quoted with the headline, report-only (house rule).
- PBO, observed OOS Sharpe, Bailey–LdP DSR reported.
One run per registration version. No parameter iteration of any kind.
8. Kill criteria (ALL must hold to survive)
Judged on the binding configuration (pinned costs, haircut applied):
- PBO ≤ 50% (no worse than the flagship's honest survfree-golden level)
- CPCV OOS observed Sharpe > 0.524 (must beat the incumbent — a new sleeve that isn't better than more-of-the-flagship has no reason to exist)
- DSR ≥ 0.665 (ditto)
- Fails any of the above, or passes only at −50% costs → KILL: the entry goes on the kill list, the data assets remain owned, and this document's verdict is final for hypothesis v1.
9. What passing yields (pinned)
Survival authorizes ONLY a proposal for a paper sleeve: operator decision on the fleet-cap raise, a mandatory pessimistic-fill twin ledger (worst-side NBBO + ADV-participation haircut; paper NBBO fills flatter illiquid names), and its own pre-registered forward gate BEFORE activation. Survival authorizes no live capital, ever, by itself.
10. Implementation checklist (build before run, test on synthetic only)
Engine extensions required, each developed against synthetic fixtures with no real-universe backtest output observed until the battery run: IWM membership loader in the engine's universe path; per-bucket cost hook; terminal-haircut handling; DTC screen consumption. The battery run happens once these merge with green suites.
<!-- SENTINEL: pre-registration v1 ends here; amendments append below this line. -->Amendment A1 (2026-07-22, pre-first-run; appended per the versioning rules)
Warmup semantics discovered during engine wiring, resolved before any run: the price frame loads from 2019-01-01 solely to feed the 252-day lookback so that positions exist from the window start (2020-01-01) rather than a year in. CPCV partitions the loaded frame, so path windows overlapping 2019 exist mechanically; the 2019 portion is lookback-immature and carries the same signal-maturity property as the flagship's own data start. The judged window remains 2020-01-01 → 2025-06-30 as frozen. Data materialization trims every series to [2019-01-01, 2025-06-30] so the engine can never see bars past the frozen window edge. No text above the sentinel is modified by this amendment.
Implementation record (2026-07-22, pre-run; appended)
§10 checklist built and merged with the flagship byte-identical (all
extensions config-gated inert): per-symbol bucket costs
(costs.model: "bucket" + compute_trading_costs(rt_bps_by_symbol=),
threaded through all five engine cost sites including kill-switch/overlay
transitions), terminal delisting haircut (data.delisting_haircut_pct,
applied once on the first forward-filled day, never at the frame edge),
eligibility screens (universe.min_price, universe.dtc_screen with
missing-SI-included semantics), IWM membership via the existing
membership_filename mechanism, config/microcap_battery.yaml
(deliberately NOT a registered sleeve — the registry contract is for
trading sleeves), and thales run-microcap-battery with mechanical one-run
enforcement + the frozen-criteria verdict.
Data addition (2026-07-22, pre-run; appended — §7 allows adding data)
Alive-gap Tiingo fetch (thales fetch-microcap-gap): 298/306 names fetched
(8 not served by Tiingo), 294 admitted through
validate_prices_for_ingestion (4 quarantined, bytes deleted) into
data/microcap/gap_prices/. Manifest updated in place: priced universe now
343 raw + 3,045 SimFin + 294 tiingo_gap = 3,682 of 6,418; residual
gap_alive = 12, gap_dead = 2,721 (pre-2020-era, outside the frozen window's
survivorship-clean claim). data/raw untouched. This is the final data
state before the run.
Attempt record (2026-07-22, appended)
The first firing of the run was KILLED ~27 minutes in (machine sleep;
multiprocessing pool killed hard) after completing materialization, the
binding CPCV, and the ×0.5 sensitivity — no verdict was rendered. Per the
one-run rule's intent (one completed JUDGMENT, no result-shopping), the
runner gained a --resume mode: allowed only while no verdict.json exists,
reusing completed variant artifacts byte-identically (deterministic
engine), never re-choosing anything. A rendered verdict remains
unrepeatable. The resumed firing runs under caffeinate so sleep cannot
kill it again.