Micro/small-cap momentum — data feasibility scoping (2026-07-21)
RESOLUTION 2026-07-22 — all three open questions answered empirically; verdict: BUILDABLE, with a pre-registered scope constraint. See "Open questions — RESOLVED" at the bottom. Headline: monthly PiT Russell 2000 membership is recoverable 2006→present from Internet-Archive-captured iShares IWM holdings (riazarbi is IVV-only; IWC is unusable — 14 scattered snapshots); EDGAR serves micro share counts to the ~2010-12 XBRL floor; SimFin prices 46% of the measured true-delist pool (2020-25 era, shares outstanding on every covered name). The binding constraint is pre-2020 dead-micro PRICES — the survivorship-clean battery window is therefore 2020+, with earlier eras admitted only under a documented survivorship caveat or after a Wayback micro-recovery pilot measures its yield.
Status: SCOPING → VERIFIED, build phase authorized to be proposed. No sleeve, no fleet-cap raise, no backtest, no config change is authorized by this memo. It defines the data program that must exist — and the pre-registrations that must be written — before the first micro-cap backtest is allowed to run. Nothing here touches the December-gated flagship.
Why this direction (adjudicated 2026-07-21)
An external strategy review argued, and the research record supports, that the one honest growth direction for this apparatus is capacity-constrained small/micro-cap cross-sectional momentum:
- McLean–Pontiff (2016): post-publication factor decay concentrates where arbitrage is cheap; residual predictability survives in high-idiosyncratic- risk, low-liquidity names — the corner institutions structurally cannot deploy in and a ~$34k book can.
- Our own kill list scopes "factor space exhausted" to large-cap (Russell 1000 on this apparatus). Micro/small-cap is genuinely outside the killed space — this is an extension, not a relitigation.
- The apparatus transfers whole: survivorship-free CPCV, golden-store ingestion gates, timing-luck sweep, pre-registration discipline, the falsification stance.
The same adjudication retired the chains/skew scarcity claim (vendor EOD
option archives exist at retail prices — our own 2026-06-08 feasibility memo
already priced ORATS at ~$400–600; the capture stays as survivorship-clean,
same-feed research convenience with its pre-registered escalation trigger
unchanged) and added the FINRA short-sale capture (registered today in
CAPTURES.md: short_volume + short_interest) whose intended consumer is
this program's borrow/crowding screen.
Standing constraints (non-negotiable)
- Battery before sleeve. A paper account exists only for (a) a strategy that survived its pre-registered battery, or (b) a designated negative control. Fleet cap 3 stands until there is a survivor.
- Pre-register before peeking. The cost model, delisting-return policy, universe definition, and kill criteria below must be pinned in config/a hash-frozen memo BEFORE the first backtest run — else this becomes cost-model/universe shopping.
- The flagship is untouchable through the December gate; this program competes only for background agent-hours.
Workstream 1 — Point-in-time universe (the crux)
Two candidate definitions, to be resolved by verification:
- A. Index-membership (iShares IWM / IWC holdings). Same mechanism as our IVV membership (riazarbi archive). OPEN QUESTION #1: verify the riazarbi archive actually covers IWM/IWC and how deep. Expected ~2019+ at best — thin for a full-history battery.
- B. Cap-rank-defined universe from owned data (preferred if buildable).
Universe = names ranked ~1001–3000 by PiT market cap, computed from owned
prices × PiT shares outstanding (SEC EDGAR
dei:EntityCommonSharesOutstanding/ SimFin share counts). No index vendor, no license, reconstructible for any date the price+shares stores cover. OPEN QUESTION #2: PiT share-count coverage for sub-$500M names, especially pre-2015.
Rule 1 (single resolution point per data-kind) applies: whichever wins
becomes the one load_microcap_universe path.
Workstream 2 — Survivorship (harder than large-cap, and decisive)
Micro-cap delisting rates are a multiple of large-cap and delisting returns are brutal; a survivor-biased micro-cap momentum backtest would be worse than no backtest. Requirements:
- Extend the golden-store stitch (Tiingo live + SimFin delisted + Wayback pre-2020) to the micro universe. OPEN QUESTION #3: measured delisted coverage below $500M — SimFin is B-grade in large-cap; quantify, don't assume, in micro.
- Pre-registered delisting-return policy (Shumway-class): where the final return is unobserved, impute a pinned conservative haircut (literature anchors ≈ −30% NYSE/AMEX, −55% Nasdaq performance delists) — the exact constants go in the battery pre-registration, chosen BEFORE any backtest.
- Every source crosses
validate_prices_for_ingestionas usual.
Workstream 3 — Costs and capacity (the honest killer)
Flat 10bps is invalid below ~$500M market cap. Requirements:
- Parameterize the existing Almgren–Chriss model (
backtest/costs.py) by ADV/spread buckets estimated from owned OHLCV; expect 30–100bp+ effective round-trips in true micro names. - Pre-register the cost model before the battery — it is part of the hypothesis, not a tuning knob. Sensitivity report (±50% cost) required in the battery output.
- Capacity ceiling via
backtest/capacity.py; the answer bounds position sizes and (later) any live allocation.
Workstream 4 — Execution realism (paper is MORE flattering here)
Alpaca paper fills at NBBO with full size at touch — in illiquid names that overstates fills exactly where it matters. Any future paper sleeve requires a pessimistic twin analogue (VRP precedent): mark every fill at worst-side NBBO plus an ADV-participation haircut; the honest track lives between the curves. Design the ledger with the battery, not after activation.
Workstream 5 — Borrow/crowding screen (consumer of today's capture)
Long-only book; the Muravyev–Pearson–Pollet borrow-fee-artifact class must be
screened OUT, never harvested. Pre-register a days-to-cover / short-interest
exclusion threshold (from data/short_interest) at the construction stage.
Screen constants pinned with the battery.
The battery (to be pre-registered in full before first run)
Same standard as the flagship, no discounts: survivorship-free CPCV on the extended golden store (PBO, DSR, IS-OOS corr), walk-forward with bootstrap CIs, start-day sweep quoted with any headline, pre-pinned cost model and delisting policy, and written kill criteria (e.g., "net-of-cost survfree observed Sharpe below the flagship's 0.524 ⇒ kill — the cost drag ate the spread"; exact criteria pinned in the battery memo).
Sequencing & effort (background agent-hours only)
- Resolve OPEN QUESTIONS #1–#3 (verification, ~1 session).
- Build universe + survivorship extension behind ingestion gates (~2–4 sessions), then the cost parameterization (~1–2 sessions).
- Write + hash-freeze the battery pre-registration.
- Run the battery. Survivor → then and only then discuss a sleeve (fleet-cap raise, its own account, pessimistic ledger). Non-survivor → the kill list gains its first micro-cap entry, and the data assets remain owned.
Open questions — RESOLVED (verification session 2026-07-22)
OQ#1 — PiT small/micro universe: ANSWERED — Option A wins, via Wayback-archived IWM
- riazarbi is IVV-only. Full repo sweep:
sp500-scraperis the only index archive (itsishares/tree contains justsp500/); no IWM/IWC anywhere in the account. - iShares-direct is still dead (the 2026-06-07 verdict re-confirmed): the holdings ajax endpoint returns HTML to non-browser requests, current and historical alike.
- BUT the Internet Archive holds the IWM holdings endpoint with historical
asOfDaterequests: 245 distinct month-end as-of dates spanning 2006-09-29 → 2026-07-10 (~12/year every year), plus 48 capture-day snapshots. Payloads verified parseable across eras (2006, 2012, 2019, 2021, 2023, 2024 samples: ~1,966–2,035 clean US-equity rows each, JSONaaDataor CSV; per-row market value included). One junk capture found (2021-12) — fetch-with-fallback across candidate captures handles it. - IWC (micro-cap ETF) is unusable as a series: 14 scattered as-of dates 2007→2026. The micro slice is therefore defined by cap-rank within/below the R2000 universe, not by IWC membership.
- Build:
build_membership_from_wayback_iwmfollowing the existingwayback.py+ IVV-builder patterns (CDX enumerate → fetch-with-fallback → clean → long format). Same license posture as IVV: facts only, output gitignored. Session artifacts were scratchpad-ephemeral; the builder re-derives from CDX.
OQ#2 — PiT micro share counts: ANSWERED — feasible with a ~2010-12 floor
- Probe: 12 representative bottom-third living 2024 IWM members → 9/12
resolve in the current SEC ticker map; all 3 misses are names that died
after the snapshot (Aaron's, Surmodics, MoneyLion) — the map is
current-only, so dead-name resolution needs a historical ticker→CIK source
(EDGAR submissions' former-ticker fields, or archived
company_tickers.jsonsnapshots; both free). - Of the resolved: share-count concepts present on effectively all;
established filers reach 2010 (XBRL phase-in floor), recent IPOs start
at their IPO (correct, not a gap). 6/9 reach ≤2013. Concept variance
(dei cover-page vs us-gaap variants, dual-class like UONE/UONEK) is the
same fallback-engineering class
fundamentals.pyalready solved for the R1000 (CONCEPT_FALLBACKS). - Dead names: SimFin carries
Shares Outstandingon 149/149 of its covered dead-pool names — price and shares arrive together. - Net: cap-ranking a PiT micro universe is high-confidence 2012+, decent 2010+, and thin before 2010 (pre-XBRL) — but prices, not shares, are the binding constraint (below).
OQ#3 — sub-$500M delisted coverage: ANSWERED — 46% out-of-box, gap targetable, pre-2020 binding
- Honest test population from real churn: 1,996 members in the 2019-12-31 IWM snapshot; 435 disappeared by 2021-12 and never returned through 2024-06. Of those: 30 graduated into our R1000 store (alive), 81 still SEC-registered today (alive/renamed — recoverable live), 324 true delist candidates.
- SimFin daily bars cover 149/324 = 46%, with plausible terminal years (2020: 16, 2021: 61, 2022: 17, 2023: 19, 2024: 12, 2025: 24) and shares outstanding on every covered name. Alive-pool control: 69/81 covered.
- Wayback's current store covers 3/324 — it was built for pre-2020/GFC S&P names. The machinery can be pointed at the 175-name gap (known one-session pilot pattern); yield unknown until piloted.
- The binding constraint: pre-2020 dead-micro prices. SimFin's delisted strength begins ~2019-20; no free source has been verified for dead micro names before that. Therefore the pre-registered scope decision for the battery: survivorship-clean window = 2020+ only (~6y and growing); any pre-2020 extension enters either (a) with a documented survivorship caveat attached to every number it produces, or (b) after a Wayback micro-recovery pilot measures actual pre-2020 yield. Choosing (a) or (b) after seeing backtest results would be scope-shopping — the choice is made at battery pre-registration, before the first run.
Build-phase order (unchanged from the memo, now de-risked)
build_membership_from_wayback_iwm(monthly R2000 PiT, 2006→present).- SimFin micro price+shares stitch behind
validate_prices_for_ingestion; optional Wayback micro pilot for the 175-name gap. - Cost parameterization (Almgren-Chriss by ADV/spread buckets) + capacity.
- Battery pre-registration (hash-frozen: universe def, cost model, delisting haircuts, DTC screen, kill criteria, 2020+ window) — then run.
Build log
Workstream 1 COMPLETE (2026-07-22, same session as the resolution):
thales build-iwm-membership shipped (constituents.py: CDX enumeration →
gzip-aware fetch with junk-capture fallback → era-tolerant parse → resumable
cache → assembler; 7 tests). Acquisition: 236/245 as-of dates recovered
(9 unrecoverable, tolerated as holes) → r2000_membership_iwm.parquet:
464,725 rows · 236 snapshots · 6,418 unique tickers · 2006-09-29 →
2025-09-11 (gitignored, facts-only posture). QC: per-snapshot counts
1,865–2,064 with zero outliers; June reconstitution signature confirmed
(median June churn 381 vs 13–16 all other months); ETSY/SMCI enter/exit on
their known index careers. Identity hazard caught live: BBBY runs
continuously to 2023-03 (bankruptcy era) then reappears once at 2025-09 as
the RELAUNCHED issuer — same ticker, same display name, different company.
Build-phase consequence (pre-registered here): membership consumers must
key identity as (symbol, era)/CIK, never symbol alone; name-based
disambiguation is proven insufficient. Next: workstream 2 (SimFin micro
stitch behind the ingestion gate).
Workstream 2 COMPLETE (2026-07-22, same session): thales build-microcap-prices shipped (data/microcap.py; golden PASS-2
conventions — adjustment-factor O/H/L, malformed bars dropped as flagged
gaps, validate_prices_for_ingestion as the gate, contamination excludes
honored; 3 tests). Real build: universe 6,418 → raw 343 · SimFin admitted
3,045 / quarantined 0 (78 MB, 3,045 per-symbol parquets) · shares panel
3.07 M rows / 2,961 symbols · gap 3,027 (SEC-alive 306 → future Tiingo
fetch; dead 2,721 → overwhelmingly pre-2020-era names, consistent with the
2020+ window). Spot checks: ROVR bars end 4 days after its take-private
membership exit; ASXC ends at its 2024-08 acquisition; SimFin left edge
~2020-07 visible. MANIFEST.json is the coverage statement the battery
pre-registration cites. Next: workstream 3 (micro cost parameterization),
then the battery pre-registration; the 306-name Tiingo fetch and the
Wayback dead-name pilot remain optional coverage upgrades.
Workstream 3 COMPLETE (2026-07-22, same session): thales estimate-micro-costs shipped (backtest/micro_costs.py; 10 tests). Two
findings, one methodological and load-bearing:
- Estimator selection by anchor: clipped Corwin-Schultz was disqualified as primary — its per-window clipping turns volatility noise into a positive floor (~60bps of fiction on liquid R1000 names). Abdi-Ranaldo (2017) replaced it (noise cancels pre-sqrt); CS kept as a reference column.
- The resolution-floor result: at n=252 daily bars, mean-then-sqrt
spread estimation has an irreducible noise floor ≈ 2·√(se(x)) — ~60bps
at large-cap vol (the R1000 anchor reads exactly this, validating the
noise model) and 100–200bps in high-vol micro buckets. EVERY bucket
is therefore resolution-limited; the table's spreads are UPPER BOUNDS
with correct liquidity ordering (medians 227 / 106 / 58 / 60 / 65 bps by
$ADV bucket; p75 403 / 186 / 113 / 98 / 108). Per-name floors are
computed and stored; buckets carry a
resolution_limitedflag.
Consequence for the battery pre-registration (recommendation, to be
pinned at workstream 4): use the p75 upper-bound spreads as the cost
inputs — conservative by construction (overstated costs yield false kills,
never false ships — the correct error direction for a falsification
apparatus) — plus the already-required ±50% cost sensitivity. Available
refinement if the battery is cost-decisive: a 756-bar window cuts floors
~24% at the price of excluding short-history dead names. eta stays 0.1
(pinned; no fill data exists to calibrate it). Artifact:
data/microcap/cost_table.json (method, floors, both estimators).
FINAL VERDICT — hypothesis v1 KILLED (battery run completed 2026-07-22)
The one-shot battery (pre-registration hash-frozen before any backtest
existed; run record in research/microcap_battery/) rendered KILL on all
three criteria: binding PBO 62.5% (bar: ≤50%), OOS observed Sharpe
−0.304 (bar: >0.524), DSR 0.027 (bar: ≥0.665). The sensitivities
remove every escape hatch: at HALF costs the Sharpe is still −0.262
(the signal itself had negative out-of-sample edge on this universe —
costs only deepened it), at 1.5× costs −0.344, and the no-haircut run is
identical to binding (−0.304 — delisting losses were not the driver). PBO
is 62.5% in every variant. The start-day sweep is negative across
schedules.
What is dead: the incumbent 252/5 momentum+smoothness signal applied to the PiT Russell 2000, 2020-01→2025-06, under honest upper-bound costs — hypothesis v1, final per the pre-registration. The McLean–Pontiff "residual edge where arbitrage is costly" thesis did NOT rescue this signal in this window. Any post-hoc regime narrative (2021 meme reversals, 2022 small-cap bear) is exactly what the pre-registration exists to preempt: the window was pinned for data honesty before any number was seen.
What remains owned: the 2006→2025 PiT membership record, the 3,682-name gate-validated price stores, the shares panel, the cost artifact + its resolution-floor methodology, the FINRA archive, and the config-gated engine extensions (bucket costs / haircut / screens — reusable by any future registration). A v2 hypothesis (different signal family or window) requires its own pre-registration. The flagship and the December gate are untouched.