# Micro-cap momentum battery — PRE-REGISTRATION v1 (2026-07-22)

**Binding document.** Everything above the sentinel line is hash-frozen by
`tests/test_docs_consistency.py` the moment this merges. Amendments are
allowed ONLY before the first battery run, as a dated superseding section
appended below the sentinel plus an explicit hash update in the test — a
recorded act, never a silent edit. After the first run, the results stand
against exactly this text. Choosing any value below after seeing micro-cap
backtest results would be the gate-shopping this document exists to prevent;
as of this writing, **zero micro-cap backtests have ever been run** on this
apparatus.

Evidence base: `research/2026-07-21_microcap_data_feasibility.md`
(resolution + build log, workstreams 1–3) and
`data/microcap/cost_table.json` (measurement artifact).

## 1. Hypothesis (exact, no signal tuning)

The flagship's **incumbent registered signal** — 252-day momentum, skip 5,
blended 50/50 with 21-day smoothness, features shifted one day — applied to
a small/micro-cap universe, survives honest (upper-bound) costs because
post-publication factor decay concentrates where arbitrage capital cannot
deploy (McLean–Pontiff). **No signal parameter may differ from the
flagship's production values.** A variant signal is a different hypothesis
requiring its own pre-registration.

## 2. Universe (pinned)

- **Membership**: point-in-time Russell 2000 via the Wayback-IWM record
  (`data/constituents/r2000_membership_iwm.parquet`); a name is a member on
  date *d* iff present in the most recent snapshot ≤ *d* (monthly cadence;
  the 9 unrecoverable snapshots are tolerated holes).
- **Identity**: (symbol, membership era) — never symbol alone (BBBY
  ticker-reuse finding, build log).
- **Eligibility at selection** (non-performance justifications only):
  price ≥ $1.00 (exchange delisting-standard threshold), ≥ 60 prior daily
  bars in the owned stores (estimator/covariance floor), and the borrow
  screen of §5.
- **Selection**: top **50** eligible members by the signal (flagship's N),
  monthly on the first trading day, daily risk management between — the
  flagship's registered cadence, unchanged.
- **Construction**: identical to the flagship's registered production
  construction (HRP weighting, half-Kelly, 10% position cap, 35% sector
  cap, 12% vol target VIX-scaled, 5% daily turnover cap
  frequency-scaled, 20% kill-switch). Universe, costs, delisting handling
  and the borrow screen are the ONLY differences from the flagship.

## 3. Window (pinned)

**2020-01-01 → 2025-06-30.** The survivorship-clean span (SimFin's delisted
coverage begins ~2019-20; feasibility memo OQ#3). Pre-2020 runs are
forbidden under this registration — extending backward requires a
superseding version with a measured Wayback micro pilot or an attached
survivorship caveat on every number.

## 4. Costs (pinned upper bounds — false kills over false ships)

Per-name flat round-trip assigned by trailing-252-bar median dollar ADV at
selection, from the measured p75 upper-bound table
(`cost_table.json`, Abdi-Ranaldo primary, resolution floors documented):

| $ADV bucket | round-trip bps (binding) |
|---|---|
| < $250k | **414.2** |
| $250k–1M | **194.2** |
| $1M–5M | **120.0** |
| $5M–25M | **103.7** |
| ≥ $25M | **114.4** |

(= p75 AR spread + Almgren-Chriss impact at 1% ADV participation, η = 0.1
pinned — uncalibratable without fill data.) Sensitivity at ±50% of this
table is REQUIRED in the battery output and is informational — **except: a
configuration that passes only under −50% costs is a KILL.**

## 5. Borrow/crowding screen (pinned)

Exclude at selection any name whose most recent FINRA consolidated
short-interest settlement (≤ selection date, `data/short_interest*`) shows
**days-to-cover > 10.0** — the Muravyev–Pearson–Pollet borrow-fee-artifact
class is screened OUT, never harvested. Names with no SI record are
INCLUDED (missing ≠ shorted).

## 6. Delisting returns (pinned)

If a HELD name's price series terminates while held (no further bars), the
final holding return takes an additional **−55%** terminal haircut (the
harsher Shumway-class Nasdaq figure, applied uniformly because our data
cannot distinguish deaths from acquisitions; over-punitive for acquisitions
— stated, and in the false-kill direction). A no-haircut run is REQUIRED in
the output as sensitivity, and is non-binding.

## 7. The battery (pinned)

On the stores as they exist at run time (microcap store + data/raw; the
optional 306-name Tiingo alive-gap fetch may be added BEFORE the run —
adding data is not tuning):

1. Survivorship-free CPCV: n_groups = 6, k_test = 2, embargo = 5,
   purge = max(252, `min_required_purge(config)`).
2. Walk-forward with bootstrap CIs.
3. Start-day sweep — quoted with the headline, report-only (house rule).
4. PBO, observed OOS Sharpe, Bailey–LdP DSR reported.

One run per registration version. No parameter iteration of any kind.

## 8. Kill criteria (ALL must hold to survive)

Judged on the binding configuration (pinned costs, haircut applied):

- **PBO ≤ 50%** (no worse than the flagship's honest survfree-golden level)
- **CPCV OOS observed Sharpe > 0.524** (must beat the incumbent — a new
  sleeve that isn't better than more-of-the-flagship has no reason to exist)
- **DSR ≥ 0.665** (ditto)
- Fails any of the above, or passes only at −50% costs → **KILL**: the
  entry goes on the kill list, the data assets remain owned, and this
  document's verdict is final for hypothesis v1.

## 9. What passing yields (pinned)

Survival authorizes ONLY a **proposal** for a paper sleeve: operator
decision on the fleet-cap raise, a mandatory pessimistic-fill twin ledger
(worst-side NBBO + ADV-participation haircut; paper NBBO fills flatter
illiquid names), and its own pre-registered forward gate BEFORE activation.
Survival authorizes no live capital, ever, by itself.

## 10. Implementation checklist (build before run, test on synthetic only)

Engine extensions required, each developed against synthetic fixtures with
no real-universe backtest output observed until the battery run: IWM
membership loader in the engine's universe path; per-bucket cost hook;
terminal-haircut handling; DTC screen consumption. The battery run happens
once these merge with green suites.

<!-- SENTINEL: pre-registration v1 ends here; amendments append below this line. -->

## Amendment A1 (2026-07-22, pre-first-run; appended per the versioning rules)

**Warmup semantics discovered during engine wiring, resolved before any
run:** the price frame loads from **2019-01-01** solely to feed the 252-day
lookback so that positions exist from the window start (2020-01-01) rather
than a year in. CPCV partitions the loaded frame, so path windows
overlapping 2019 exist mechanically; the 2019 portion is lookback-immature
and carries the same signal-maturity property as the flagship's own data
start. The judged window remains 2020-01-01 → 2025-06-30 as frozen. Data
materialization trims every series to [2019-01-01, 2025-06-30] so the
engine can never see bars past the frozen window edge. No text above the
sentinel is modified by this amendment.

## Implementation record (2026-07-22, pre-run; appended)

§10 checklist built and merged with the flagship byte-identical (all
extensions config-gated inert): per-symbol bucket costs
(`costs.model: "bucket"` + `compute_trading_costs(rt_bps_by_symbol=)`,
threaded through all five engine cost sites including kill-switch/overlay
transitions), terminal delisting haircut (`data.delisting_haircut_pct`,
applied once on the first forward-filled day, never at the frame edge),
eligibility screens (`universe.min_price`, `universe.dtc_screen` with
missing-SI-included semantics), IWM membership via the existing
`membership_filename` mechanism, `config/microcap_battery.yaml`
(deliberately NOT a registered sleeve — the registry contract is for
trading sleeves), and `thales run-microcap-battery` with mechanical one-run
enforcement + the frozen-criteria verdict.

## Data addition (2026-07-22, pre-run; appended — §7 allows adding data)

Alive-gap Tiingo fetch (`thales fetch-microcap-gap`): 298/306 names fetched
(8 not served by Tiingo), **294 admitted** through
`validate_prices_for_ingestion` (4 quarantined, bytes deleted) into
`data/microcap/gap_prices/`. Manifest updated in place: priced universe now
**343 raw + 3,045 SimFin + 294 tiingo_gap = 3,682 of 6,418**; residual
gap_alive = 12, gap_dead = 2,721 (pre-2020-era, outside the frozen window's
survivorship-clean claim). `data/raw` untouched. This is the final data
state before the run.

## Attempt record (2026-07-22, appended)

The first firing of the run was KILLED ~27 minutes in (machine sleep;
multiprocessing pool killed hard) after completing materialization, the
binding CPCV, and the ×0.5 sensitivity — no verdict was rendered. Per the
one-run rule's intent (one completed JUDGMENT, no result-shopping), the
runner gained a `--resume` mode: allowed only while no verdict.json exists,
reusing completed variant artifacts byte-identically (deterministic
engine), never re-choosing anything. A rendered verdict remains
unrepeatable. The resumed firing runs under `caffeinate` so sleep cannot
kill it again.
