# Micro/small-cap momentum — data feasibility scoping (2026-07-21)

> **RESOLUTION 2026-07-22 — all three open questions answered empirically;
> verdict: BUILDABLE, with a pre-registered scope constraint.** See "Open
> questions — RESOLVED" at the bottom. Headline: monthly PiT Russell 2000
> membership is recoverable 2006→present from Internet-Archive-captured
> iShares IWM holdings (riazarbi is IVV-only; IWC is unusable — 14 scattered
> snapshots); EDGAR serves micro share counts to the ~2010-12 XBRL floor;
> SimFin prices 46% of the measured true-delist pool (2020-25 era, shares
> outstanding on every covered name). **The binding constraint is pre-2020
> dead-micro PRICES** — the survivorship-clean battery window is therefore
> **2020+**, with earlier eras admitted only under a documented survivorship
> caveat or after a Wayback micro-recovery pilot measures its yield.

**Status: SCOPING → VERIFIED, build phase authorized to be proposed.** No
sleeve, no fleet-cap raise, no backtest, no config change is authorized by
this memo. It defines the data program that must
exist — and the pre-registrations that must be written — *before* the first
micro-cap backtest is allowed to run. Nothing here touches the December-gated
flagship.

## Why this direction (adjudicated 2026-07-21)

An external strategy review argued, and the research record supports, that
the one honest growth direction for this apparatus is capacity-constrained
small/micro-cap cross-sectional momentum:

- McLean–Pontiff (2016): post-publication factor decay concentrates where
  arbitrage is cheap; residual predictability survives in high-idiosyncratic-
  risk, low-liquidity names — the corner institutions structurally cannot
  deploy in and a ~$34k book can.
- Our own kill list scopes "factor space exhausted" to **large-cap** (Russell
  1000 on this apparatus). Micro/small-cap is genuinely outside the killed
  space — this is an extension, not a relitigation.
- The apparatus transfers whole: survivorship-free CPCV, golden-store
  ingestion gates, timing-luck sweep, pre-registration discipline, the
  falsification stance.

The same adjudication **retired the chains/skew scarcity claim** (vendor EOD
option archives exist at retail prices — our own 2026-06-08 feasibility memo
already priced ORATS at ~$400–600; the capture stays as survivorship-clean,
same-feed research convenience with its pre-registered escalation trigger
unchanged) and **added the FINRA short-sale capture** (registered today in
CAPTURES.md: `short_volume` + `short_interest`) whose intended consumer is
this program's borrow/crowding screen.

## Standing constraints (non-negotiable)

1. **Battery before sleeve.** A paper account exists only for (a) a strategy
   that survived its pre-registered battery, or (b) a designated negative
   control. Fleet cap 3 stands until there is a survivor.
2. **Pre-register before peeking.** The cost model, delisting-return policy,
   universe definition, and kill criteria below must be pinned in config/a
   hash-frozen memo BEFORE the first backtest run — else this becomes
   cost-model/universe shopping.
3. **The flagship is untouchable** through the December gate; this program
   competes only for background agent-hours.

## Workstream 1 — Point-in-time universe (the crux)

Two candidate definitions, to be resolved by verification:

- **A. Index-membership (iShares IWM / IWC holdings).** Same mechanism as our
  IVV membership (riazarbi archive). **OPEN QUESTION #1: verify the riazarbi
  archive actually covers IWM/IWC and how deep.** Expected ~2019+ at best —
  thin for a full-history battery.
- **B. Cap-rank-defined universe from owned data (preferred if buildable).**
  Universe = names ranked ~1001–3000 by PiT market cap, computed from owned
  prices × PiT shares outstanding (SEC EDGAR `dei:EntityCommonSharesOutstanding`
  / SimFin share counts). No index vendor, no license, reconstructible for
  any date the price+shares stores cover. **OPEN QUESTION #2: PiT share-count
  coverage for sub-$500M names, especially pre-2015.**

Rule 1 (single resolution point per data-kind) applies: whichever wins
becomes the one `load_microcap_universe` path.

## Workstream 2 — Survivorship (harder than large-cap, and decisive)

Micro-cap delisting rates are a multiple of large-cap and delisting returns
are brutal; a survivor-biased micro-cap momentum backtest would be *worse
than no backtest*. Requirements:

- Extend the golden-store stitch (Tiingo live + SimFin delisted + Wayback
  pre-2020) to the micro universe. **OPEN QUESTION #3: measured delisted
  coverage below $500M** — SimFin is B-grade in large-cap; quantify, don't
  assume, in micro.
- **Pre-registered delisting-return policy** (Shumway-class): where the final
  return is unobserved, impute a pinned conservative haircut (literature
  anchors ≈ −30% NYSE/AMEX, −55% Nasdaq performance delists) — the exact
  constants go in the battery pre-registration, chosen BEFORE any backtest.
- Every source crosses `validate_prices_for_ingestion` as usual.

## Workstream 3 — Costs and capacity (the honest killer)

Flat 10bps is invalid below ~$500M market cap. Requirements:

- Parameterize the existing Almgren–Chriss model (`backtest/costs.py`) by
  ADV/spread buckets estimated from owned OHLCV; expect 30–100bp+ effective
  round-trips in true micro names.
- **Pre-register the cost model before the battery** — it is part of the
  hypothesis, not a tuning knob. Sensitivity report (±50% cost) required in
  the battery output.
- Capacity ceiling via `backtest/capacity.py`; the answer bounds position
  sizes and (later) any live allocation.

## Workstream 4 — Execution realism (paper is MORE flattering here)

Alpaca paper fills at NBBO with full size at touch — in illiquid names that
overstates fills exactly where it matters. Any future paper sleeve requires a
**pessimistic twin** analogue (VRP precedent): mark every fill at worst-side
NBBO plus an ADV-participation haircut; the honest track lives between the
curves. Design the ledger with the battery, not after activation.

## Workstream 5 — Borrow/crowding screen (consumer of today's capture)

Long-only book; the Muravyev–Pearson–Pollet borrow-fee-artifact class must be
screened OUT, never harvested. Pre-register a days-to-cover / short-interest
exclusion threshold (from `data/short_interest`) at the construction stage.
Screen constants pinned with the battery.

## The battery (to be pre-registered in full before first run)

Same standard as the flagship, no discounts: survivorship-free CPCV on the
extended golden store (PBO, DSR, IS-OOS corr), walk-forward with bootstrap
CIs, start-day sweep quoted with any headline, pre-pinned cost model and
delisting policy, and **written kill criteria** (e.g., "net-of-cost survfree
observed Sharpe below the flagship's 0.524 ⇒ kill — the cost drag ate the
spread"; exact criteria pinned in the battery memo).

## Sequencing & effort (background agent-hours only)

1. Resolve OPEN QUESTIONS #1–#3 (verification, ~1 session).
2. Build universe + survivorship extension behind ingestion gates (~2–4
   sessions), then the cost parameterization (~1–2 sessions).
3. Write + hash-freeze the battery pre-registration.
4. Run the battery. Survivor → then and only then discuss a sleeve (fleet-cap
   raise, its own account, pessimistic ledger). Non-survivor → the kill list
   gains its first micro-cap entry, and the data assets remain owned.

---

## Open questions — RESOLVED (verification session 2026-07-22)

### OQ#1 — PiT small/micro universe: **ANSWERED — Option A wins, via Wayback-archived IWM**

- **riazarbi is IVV-only.** Full repo sweep: `sp500-scraper` is the only
  index archive (its `ishares/` tree contains just `sp500/`); no IWM/IWC
  anywhere in the account.
- **iShares-direct is still dead** (the 2026-06-07 verdict re-confirmed):
  the holdings ajax endpoint returns HTML to non-browser requests, current
  and historical alike.
- **BUT the Internet Archive holds the IWM holdings endpoint with historical
  `asOfDate` requests: 245 distinct month-end as-of dates spanning
  2006-09-29 → 2026-07-10** (~12/year every year), plus 48 capture-day
  snapshots. Payloads verified parseable across eras (2006, 2012, 2019,
  2021, 2023, 2024 samples: ~1,966–2,035 clean US-equity rows each, JSON
  `aaData` or CSV; per-row market value included). One junk capture found
  (2021-12) — fetch-with-fallback across candidate captures handles it.
- **IWC (micro-cap ETF) is unusable as a series**: 14 scattered as-of dates
  2007→2026. The micro slice is therefore defined by **cap-rank within/below
  the R2000 universe**, not by IWC membership.
- Build: `build_membership_from_wayback_iwm` following the existing
  `wayback.py` + IVV-builder patterns (CDX enumerate → fetch-with-fallback →
  clean → long format). Same license posture as IVV: facts only, output
  gitignored. Session artifacts were scratchpad-ephemeral; the builder
  re-derives from CDX.

### OQ#2 — PiT micro share counts: **ANSWERED — feasible with a ~2010-12 floor**

- Probe: 12 representative bottom-third living 2024 IWM members → 9/12
  resolve in the current SEC ticker map; **all 3 misses are names that died
  after the snapshot** (Aaron's, Surmodics, MoneyLion) — the map is
  current-only, so dead-name resolution needs a historical ticker→CIK source
  (EDGAR submissions' former-ticker fields, or archived
  `company_tickers.json` snapshots; both free).
- Of the resolved: share-count concepts present on effectively all;
  established filers reach **2010** (XBRL phase-in floor), recent IPOs start
  at their IPO (correct, not a gap). 6/9 reach ≤2013. Concept variance
  (dei cover-page vs us-gaap variants, dual-class like UONE/UONEK) is the
  same fallback-engineering class `fundamentals.py` already solved for the
  R1000 (`CONCEPT_FALLBACKS`).
- **Dead names: SimFin carries `Shares Outstanding` on 149/149 of its
  covered dead-pool names** — price and shares arrive together.
- Net: cap-ranking a PiT micro universe is high-confidence **2012+**, decent
  **2010+**, and thin before 2010 (pre-XBRL) — but prices, not shares, are
  the binding constraint (below).

### OQ#3 — sub-$500M delisted coverage: **ANSWERED — 46% out-of-box, gap targetable, pre-2020 binding**

- **Honest test population from real churn**: 1,996 members in the 2019-12-31
  IWM snapshot; 435 disappeared by 2021-12 and never returned through
  2024-06. Of those: 30 graduated into our R1000 store (alive), 81 still
  SEC-registered today (alive/renamed — recoverable live), **324 true
  delist candidates**.
- **SimFin daily bars cover 149/324 = 46%**, with plausible terminal years
  (2020: 16, 2021: 61, 2022: 17, 2023: 19, 2024: 12, 2025: 24) and shares
  outstanding on every covered name. Alive-pool control: 69/81 covered.
- **Wayback's current store covers 3/324** — it was built for pre-2020/GFC
  S&P names. The machinery can be pointed at the 175-name gap (known
  one-session pilot pattern); yield unknown until piloted.
- **The binding constraint: pre-2020 dead-micro prices.** SimFin's delisted
  strength begins ~2019-20; no free source has been verified for dead micro
  names before that. Therefore the **pre-registered scope decision for the
  battery**: survivorship-clean window = **2020+ only** (~6y and growing);
  any pre-2020 extension enters either (a) with a documented survivorship
  caveat attached to every number it produces, or (b) after a Wayback
  micro-recovery pilot measures actual pre-2020 yield. Choosing (a) or (b)
  after seeing backtest results would be scope-shopping — the choice is made
  at battery pre-registration, before the first run.

### Build-phase order (unchanged from the memo, now de-risked)

1. `build_membership_from_wayback_iwm` (monthly R2000 PiT, 2006→present).
2. SimFin micro price+shares stitch behind `validate_prices_for_ingestion`;
   optional Wayback micro pilot for the 175-name gap.
3. Cost parameterization (Almgren-Chriss by ADV/spread buckets) + capacity.
4. Battery pre-registration (hash-frozen: universe def, cost model,
   delisting haircuts, DTC screen, kill criteria, 2020+ window) — then run.

---

## Build log

**Workstream 1 COMPLETE (2026-07-22, same session as the resolution):**
`thales build-iwm-membership` shipped (`constituents.py`: CDX enumeration →
gzip-aware fetch with junk-capture fallback → era-tolerant parse → resumable
cache → assembler; 7 tests). Acquisition: **236/245 as-of dates recovered**
(9 unrecoverable, tolerated as holes) → `r2000_membership_iwm.parquet`:
**464,725 rows · 236 snapshots · 6,418 unique tickers · 2006-09-29 →
2025-09-11** (gitignored, facts-only posture). QC: per-snapshot counts
1,865–2,064 with zero outliers; **June reconstitution signature confirmed**
(median June churn 381 vs 13–16 all other months); ETSY/SMCI enter/exit on
their known index careers. **Identity hazard caught live**: BBBY runs
continuously to 2023-03 (bankruptcy era) then reappears once at 2025-09 as
the RELAUNCHED issuer — same ticker, same display name, different company.
Build-phase consequence (pre-registered here): membership consumers must
key identity as (symbol, era)/CIK, never symbol alone; name-based
disambiguation is proven insufficient. Next: workstream 2 (SimFin micro
stitch behind the ingestion gate).

**Workstream 2 COMPLETE (2026-07-22, same session):** `thales
build-microcap-prices` shipped (`data/microcap.py`; golden PASS-2
conventions — adjustment-factor O/H/L, malformed bars dropped as flagged
gaps, `validate_prices_for_ingestion` as the gate, contamination excludes
honored; 3 tests). Real build: universe 6,418 → **raw 343 · SimFin admitted
3,045 / quarantined 0 (78 MB, 3,045 per-symbol parquets) · shares panel
3.07 M rows / 2,961 symbols · gap 3,027 (SEC-alive 306 → future Tiingo
fetch; dead 2,721 → overwhelmingly pre-2020-era names, consistent with the
2020+ window)**. Spot checks: ROVR bars end 4 days after its take-private
membership exit; ASXC ends at its 2024-08 acquisition; SimFin left edge
~2020-07 visible. MANIFEST.json is the coverage statement the battery
pre-registration cites. Next: workstream 3 (micro cost parameterization),
then the battery pre-registration; the 306-name Tiingo fetch and the
Wayback dead-name pilot remain optional coverage upgrades.

**Workstream 3 COMPLETE (2026-07-22, same session):** `thales
estimate-micro-costs` shipped (`backtest/micro_costs.py`; 10 tests). Two
findings, one methodological and load-bearing:

1. **Estimator selection by anchor**: clipped Corwin-Schultz was
   disqualified as primary — its per-window clipping turns volatility noise
   into a positive floor (~60bps of fiction on liquid R1000 names).
   Abdi-Ranaldo (2017) replaced it (noise cancels pre-sqrt); CS kept as a
   reference column.
2. **The resolution-floor result**: at n=252 daily bars, mean-then-sqrt
   spread estimation has an irreducible noise floor ≈ 2·√(se(x)) — ~60bps
   at large-cap vol (the R1000 anchor reads exactly this, validating the
   noise model) and **100–200bps in high-vol micro buckets**. EVERY bucket
   is therefore resolution-limited; the table's spreads are UPPER BOUNDS
   with correct liquidity ordering (medians 227 / 106 / 58 / 60 / 65 bps by
   $ADV bucket; p75 403 / 186 / 113 / 98 / 108). Per-name floors are
   computed and stored; buckets carry a `resolution_limited` flag.

**Consequence for the battery pre-registration (recommendation, to be
pinned at workstream 4):** use the p75 upper-bound spreads as the cost
inputs — conservative by construction (overstated costs yield false kills,
never false ships — the correct error direction for a falsification
apparatus) — plus the already-required ±50% cost sensitivity. Available
refinement if the battery is cost-decisive: a 756-bar window cuts floors
~24% at the price of excluding short-history dead names. eta stays 0.1
(pinned; no fill data exists to calibrate it). Artifact:
`data/microcap/cost_table.json` (method, floors, both estimators).

---

## FINAL VERDICT — hypothesis v1 KILLED (battery run completed 2026-07-22)

The one-shot battery (pre-registration hash-frozen before any backtest
existed; run record in `research/microcap_battery/`) rendered **KILL on all
three criteria**: binding PBO **62.5%** (bar: ≤50%), OOS observed Sharpe
**−0.304** (bar: >0.524), DSR **0.027** (bar: ≥0.665). The sensitivities
remove every escape hatch: at HALF costs the Sharpe is still **−0.262**
(the signal itself had negative out-of-sample edge on this universe —
costs only deepened it), at 1.5× costs −0.344, and the no-haircut run is
identical to binding (−0.304 — delisting losses were not the driver). PBO
is 62.5% in every variant. The start-day sweep is negative across
schedules.

**What is dead**: the incumbent 252/5 momentum+smoothness signal applied to
the PiT Russell 2000, 2020-01→2025-06, under honest upper-bound costs —
hypothesis v1, final per the pre-registration. The McLean–Pontiff
"residual edge where arbitrage is costly" thesis did NOT rescue this
signal in this window. Any post-hoc regime narrative (2021 meme reversals,
2022 small-cap bear) is exactly what the pre-registration exists to
preempt: the window was pinned for data honesty before any number was
seen.

**What remains owned**: the 2006→2025 PiT membership record, the 3,682-name
gate-validated price stores, the shares panel, the cost artifact + its
resolution-floor methodology, the FINRA archive, and the config-gated
engine extensions (bucket costs / haircut / screens — reusable by any
future registration). A v2 hypothesis (different signal family or window)
requires its own pre-registration. The flagship and the December gate are
untouched.
