Thales
← research journal
Jul 21, 2026raw markdown ↗

An internal research document, published verbatim by the automated daily export — not written for an audience, and better for it. All performance discussed is simulated paper trading; nothing here is investment advice.

Micro/small-cap momentum — data feasibility scoping (2026-07-21)

RESOLUTION 2026-07-22 — all three open questions answered empirically; verdict: BUILDABLE, with a pre-registered scope constraint. See "Open questions — RESOLVED" at the bottom. Headline: monthly PiT Russell 2000 membership is recoverable 2006→present from Internet-Archive-captured iShares IWM holdings (riazarbi is IVV-only; IWC is unusable — 14 scattered snapshots); EDGAR serves micro share counts to the ~2010-12 XBRL floor; SimFin prices 46% of the measured true-delist pool (2020-25 era, shares outstanding on every covered name). The binding constraint is pre-2020 dead-micro PRICES — the survivorship-clean battery window is therefore 2020+, with earlier eras admitted only under a documented survivorship caveat or after a Wayback micro-recovery pilot measures its yield.

Status: SCOPING → VERIFIED, build phase authorized to be proposed. No sleeve, no fleet-cap raise, no backtest, no config change is authorized by this memo. It defines the data program that must exist — and the pre-registrations that must be written — before the first micro-cap backtest is allowed to run. Nothing here touches the December-gated flagship.

Why this direction (adjudicated 2026-07-21)

An external strategy review argued, and the research record supports, that the one honest growth direction for this apparatus is capacity-constrained small/micro-cap cross-sectional momentum:

  • McLean–Pontiff (2016): post-publication factor decay concentrates where arbitrage is cheap; residual predictability survives in high-idiosyncratic- risk, low-liquidity names — the corner institutions structurally cannot deploy in and a ~$34k book can.
  • Our own kill list scopes "factor space exhausted" to large-cap (Russell 1000 on this apparatus). Micro/small-cap is genuinely outside the killed space — this is an extension, not a relitigation.
  • The apparatus transfers whole: survivorship-free CPCV, golden-store ingestion gates, timing-luck sweep, pre-registration discipline, the falsification stance.

The same adjudication retired the chains/skew scarcity claim (vendor EOD option archives exist at retail prices — our own 2026-06-08 feasibility memo already priced ORATS at ~$400–600; the capture stays as survivorship-clean, same-feed research convenience with its pre-registered escalation trigger unchanged) and added the FINRA short-sale capture (registered today in CAPTURES.md: short_volume + short_interest) whose intended consumer is this program's borrow/crowding screen.

Standing constraints (non-negotiable)

  1. Battery before sleeve. A paper account exists only for (a) a strategy that survived its pre-registered battery, or (b) a designated negative control. Fleet cap 3 stands until there is a survivor.
  2. Pre-register before peeking. The cost model, delisting-return policy, universe definition, and kill criteria below must be pinned in config/a hash-frozen memo BEFORE the first backtest run — else this becomes cost-model/universe shopping.
  3. The flagship is untouchable through the December gate; this program competes only for background agent-hours.

Workstream 1 — Point-in-time universe (the crux)

Two candidate definitions, to be resolved by verification:

  • A. Index-membership (iShares IWM / IWC holdings). Same mechanism as our IVV membership (riazarbi archive). OPEN QUESTION #1: verify the riazarbi archive actually covers IWM/IWC and how deep. Expected ~2019+ at best — thin for a full-history battery.
  • B. Cap-rank-defined universe from owned data (preferred if buildable). Universe = names ranked ~1001–3000 by PiT market cap, computed from owned prices × PiT shares outstanding (SEC EDGAR dei:EntityCommonSharesOutstanding / SimFin share counts). No index vendor, no license, reconstructible for any date the price+shares stores cover. OPEN QUESTION #2: PiT share-count coverage for sub-$500M names, especially pre-2015.

Rule 1 (single resolution point per data-kind) applies: whichever wins becomes the one load_microcap_universe path.

Workstream 2 — Survivorship (harder than large-cap, and decisive)

Micro-cap delisting rates are a multiple of large-cap and delisting returns are brutal; a survivor-biased micro-cap momentum backtest would be worse than no backtest. Requirements:

  • Extend the golden-store stitch (Tiingo live + SimFin delisted + Wayback pre-2020) to the micro universe. OPEN QUESTION #3: measured delisted coverage below $500M — SimFin is B-grade in large-cap; quantify, don't assume, in micro.
  • Pre-registered delisting-return policy (Shumway-class): where the final return is unobserved, impute a pinned conservative haircut (literature anchors ≈ −30% NYSE/AMEX, −55% Nasdaq performance delists) — the exact constants go in the battery pre-registration, chosen BEFORE any backtest.
  • Every source crosses validate_prices_for_ingestion as usual.

Workstream 3 — Costs and capacity (the honest killer)

Flat 10bps is invalid below ~$500M market cap. Requirements:

  • Parameterize the existing Almgren–Chriss model (backtest/costs.py) by ADV/spread buckets estimated from owned OHLCV; expect 30–100bp+ effective round-trips in true micro names.
  • Pre-register the cost model before the battery — it is part of the hypothesis, not a tuning knob. Sensitivity report (±50% cost) required in the battery output.
  • Capacity ceiling via backtest/capacity.py; the answer bounds position sizes and (later) any live allocation.

Workstream 4 — Execution realism (paper is MORE flattering here)

Alpaca paper fills at NBBO with full size at touch — in illiquid names that overstates fills exactly where it matters. Any future paper sleeve requires a pessimistic twin analogue (VRP precedent): mark every fill at worst-side NBBO plus an ADV-participation haircut; the honest track lives between the curves. Design the ledger with the battery, not after activation.

Workstream 5 — Borrow/crowding screen (consumer of today's capture)

Long-only book; the Muravyev–Pearson–Pollet borrow-fee-artifact class must be screened OUT, never harvested. Pre-register a days-to-cover / short-interest exclusion threshold (from data/short_interest) at the construction stage. Screen constants pinned with the battery.

The battery (to be pre-registered in full before first run)

Same standard as the flagship, no discounts: survivorship-free CPCV on the extended golden store (PBO, DSR, IS-OOS corr), walk-forward with bootstrap CIs, start-day sweep quoted with any headline, pre-pinned cost model and delisting policy, and written kill criteria (e.g., "net-of-cost survfree observed Sharpe below the flagship's 0.524 ⇒ kill — the cost drag ate the spread"; exact criteria pinned in the battery memo).

Sequencing & effort (background agent-hours only)

  1. Resolve OPEN QUESTIONS #1–#3 (verification, ~1 session).
  2. Build universe + survivorship extension behind ingestion gates (~2–4 sessions), then the cost parameterization (~1–2 sessions).
  3. Write + hash-freeze the battery pre-registration.
  4. Run the battery. Survivor → then and only then discuss a sleeve (fleet-cap raise, its own account, pessimistic ledger). Non-survivor → the kill list gains its first micro-cap entry, and the data assets remain owned.

Open questions — RESOLVED (verification session 2026-07-22)

OQ#1 — PiT small/micro universe: ANSWERED — Option A wins, via Wayback-archived IWM

  • riazarbi is IVV-only. Full repo sweep: sp500-scraper is the only index archive (its ishares/ tree contains just sp500/); no IWM/IWC anywhere in the account.
  • iShares-direct is still dead (the 2026-06-07 verdict re-confirmed): the holdings ajax endpoint returns HTML to non-browser requests, current and historical alike.
  • BUT the Internet Archive holds the IWM holdings endpoint with historical asOfDate requests: 245 distinct month-end as-of dates spanning 2006-09-29 → 2026-07-10 (~12/year every year), plus 48 capture-day snapshots. Payloads verified parseable across eras (2006, 2012, 2019, 2021, 2023, 2024 samples: ~1,966–2,035 clean US-equity rows each, JSON aaData or CSV; per-row market value included). One junk capture found (2021-12) — fetch-with-fallback across candidate captures handles it.
  • IWC (micro-cap ETF) is unusable as a series: 14 scattered as-of dates 2007→2026. The micro slice is therefore defined by cap-rank within/below the R2000 universe, not by IWC membership.
  • Build: build_membership_from_wayback_iwm following the existing wayback.py + IVV-builder patterns (CDX enumerate → fetch-with-fallback → clean → long format). Same license posture as IVV: facts only, output gitignored. Session artifacts were scratchpad-ephemeral; the builder re-derives from CDX.

OQ#2 — PiT micro share counts: ANSWERED — feasible with a ~2010-12 floor

  • Probe: 12 representative bottom-third living 2024 IWM members → 9/12 resolve in the current SEC ticker map; all 3 misses are names that died after the snapshot (Aaron's, Surmodics, MoneyLion) — the map is current-only, so dead-name resolution needs a historical ticker→CIK source (EDGAR submissions' former-ticker fields, or archived company_tickers.json snapshots; both free).
  • Of the resolved: share-count concepts present on effectively all; established filers reach 2010 (XBRL phase-in floor), recent IPOs start at their IPO (correct, not a gap). 6/9 reach ≤2013. Concept variance (dei cover-page vs us-gaap variants, dual-class like UONE/UONEK) is the same fallback-engineering class fundamentals.py already solved for the R1000 (CONCEPT_FALLBACKS).
  • Dead names: SimFin carries Shares Outstanding on 149/149 of its covered dead-pool names — price and shares arrive together.
  • Net: cap-ranking a PiT micro universe is high-confidence 2012+, decent 2010+, and thin before 2010 (pre-XBRL) — but prices, not shares, are the binding constraint (below).

OQ#3 — sub-$500M delisted coverage: ANSWERED — 46% out-of-box, gap targetable, pre-2020 binding

  • Honest test population from real churn: 1,996 members in the 2019-12-31 IWM snapshot; 435 disappeared by 2021-12 and never returned through 2024-06. Of those: 30 graduated into our R1000 store (alive), 81 still SEC-registered today (alive/renamed — recoverable live), 324 true delist candidates.
  • SimFin daily bars cover 149/324 = 46%, with plausible terminal years (2020: 16, 2021: 61, 2022: 17, 2023: 19, 2024: 12, 2025: 24) and shares outstanding on every covered name. Alive-pool control: 69/81 covered.
  • Wayback's current store covers 3/324 — it was built for pre-2020/GFC S&P names. The machinery can be pointed at the 175-name gap (known one-session pilot pattern); yield unknown until piloted.
  • The binding constraint: pre-2020 dead-micro prices. SimFin's delisted strength begins ~2019-20; no free source has been verified for dead micro names before that. Therefore the pre-registered scope decision for the battery: survivorship-clean window = 2020+ only (~6y and growing); any pre-2020 extension enters either (a) with a documented survivorship caveat attached to every number it produces, or (b) after a Wayback micro-recovery pilot measures actual pre-2020 yield. Choosing (a) or (b) after seeing backtest results would be scope-shopping — the choice is made at battery pre-registration, before the first run.

Build-phase order (unchanged from the memo, now de-risked)

  1. build_membership_from_wayback_iwm (monthly R2000 PiT, 2006→present).
  2. SimFin micro price+shares stitch behind validate_prices_for_ingestion; optional Wayback micro pilot for the 175-name gap.
  3. Cost parameterization (Almgren-Chriss by ADV/spread buckets) + capacity.
  4. Battery pre-registration (hash-frozen: universe def, cost model, delisting haircuts, DTC screen, kill criteria, 2020+ window) — then run.

Build log

Workstream 1 COMPLETE (2026-07-22, same session as the resolution): thales build-iwm-membership shipped (constituents.py: CDX enumeration → gzip-aware fetch with junk-capture fallback → era-tolerant parse → resumable cache → assembler; 7 tests). Acquisition: 236/245 as-of dates recovered (9 unrecoverable, tolerated as holes) → r2000_membership_iwm.parquet: 464,725 rows · 236 snapshots · 6,418 unique tickers · 2006-09-29 → 2025-09-11 (gitignored, facts-only posture). QC: per-snapshot counts 1,865–2,064 with zero outliers; June reconstitution signature confirmed (median June churn 381 vs 13–16 all other months); ETSY/SMCI enter/exit on their known index careers. Identity hazard caught live: BBBY runs continuously to 2023-03 (bankruptcy era) then reappears once at 2025-09 as the RELAUNCHED issuer — same ticker, same display name, different company. Build-phase consequence (pre-registered here): membership consumers must key identity as (symbol, era)/CIK, never symbol alone; name-based disambiguation is proven insufficient. Next: workstream 2 (SimFin micro stitch behind the ingestion gate).

Workstream 2 COMPLETE (2026-07-22, same session): thales build-microcap-prices shipped (data/microcap.py; golden PASS-2 conventions — adjustment-factor O/H/L, malformed bars dropped as flagged gaps, validate_prices_for_ingestion as the gate, contamination excludes honored; 3 tests). Real build: universe 6,418 → raw 343 · SimFin admitted 3,045 / quarantined 0 (78 MB, 3,045 per-symbol parquets) · shares panel 3.07 M rows / 2,961 symbols · gap 3,027 (SEC-alive 306 → future Tiingo fetch; dead 2,721 → overwhelmingly pre-2020-era names, consistent with the 2020+ window). Spot checks: ROVR bars end 4 days after its take-private membership exit; ASXC ends at its 2024-08 acquisition; SimFin left edge ~2020-07 visible. MANIFEST.json is the coverage statement the battery pre-registration cites. Next: workstream 3 (micro cost parameterization), then the battery pre-registration; the 306-name Tiingo fetch and the Wayback dead-name pilot remain optional coverage upgrades.

Workstream 3 COMPLETE (2026-07-22, same session): thales estimate-micro-costs shipped (backtest/micro_costs.py; 10 tests). Two findings, one methodological and load-bearing:

  1. Estimator selection by anchor: clipped Corwin-Schultz was disqualified as primary — its per-window clipping turns volatility noise into a positive floor (~60bps of fiction on liquid R1000 names). Abdi-Ranaldo (2017) replaced it (noise cancels pre-sqrt); CS kept as a reference column.
  2. The resolution-floor result: at n=252 daily bars, mean-then-sqrt spread estimation has an irreducible noise floor ≈ 2·√(se(x)) — ~60bps at large-cap vol (the R1000 anchor reads exactly this, validating the noise model) and 100–200bps in high-vol micro buckets. EVERY bucket is therefore resolution-limited; the table's spreads are UPPER BOUNDS with correct liquidity ordering (medians 227 / 106 / 58 / 60 / 65 bps by $ADV bucket; p75 403 / 186 / 113 / 98 / 108). Per-name floors are computed and stored; buckets carry a resolution_limited flag.

Consequence for the battery pre-registration (recommendation, to be pinned at workstream 4): use the p75 upper-bound spreads as the cost inputs — conservative by construction (overstated costs yield false kills, never false ships — the correct error direction for a falsification apparatus) — plus the already-required ±50% cost sensitivity. Available refinement if the battery is cost-decisive: a 756-bar window cuts floors ~24% at the price of excluding short-history dead names. eta stays 0.1 (pinned; no fill data exists to calibrate it). Artifact: data/microcap/cost_table.json (method, floors, both estimators).


FINAL VERDICT — hypothesis v1 KILLED (battery run completed 2026-07-22)

The one-shot battery (pre-registration hash-frozen before any backtest existed; run record in research/microcap_battery/) rendered KILL on all three criteria: binding PBO 62.5% (bar: ≤50%), OOS observed Sharpe −0.304 (bar: >0.524), DSR 0.027 (bar: ≥0.665). The sensitivities remove every escape hatch: at HALF costs the Sharpe is still −0.262 (the signal itself had negative out-of-sample edge on this universe — costs only deepened it), at 1.5× costs −0.344, and the no-haircut run is identical to binding (−0.304 — delisting losses were not the driver). PBO is 62.5% in every variant. The start-day sweep is negative across schedules.

What is dead: the incumbent 252/5 momentum+smoothness signal applied to the PiT Russell 2000, 2020-01→2025-06, under honest upper-bound costs — hypothesis v1, final per the pre-registration. The McLean–Pontiff "residual edge where arbitrage is costly" thesis did NOT rescue this signal in this window. Any post-hoc regime narrative (2021 meme reversals, 2022 small-cap bear) is exactly what the pre-registration exists to preempt: the window was pinned for data honesty before any number was seen.

What remains owned: the 2006→2025 PiT membership record, the 3,682-name gate-validated price stores, the shares panel, the cost artifact + its resolution-floor methodology, the FINRA archive, and the config-gated engine extensions (bucket costs / haircut / screens — reusable by any future registration). A v2 hypothesis (different signal family or window) requires its own pre-registration. The flagship and the December gate are untouched.