Thales
methodology & verdicts

Research, honestly

Most quantitative trading research you can find online shows you its winners. This page works the other way: it lists the rules that make results trustworthy, and then every idea we built that failed — because a research process that cannot show you its dead ends is indistinguishable from luck.

The backtest is a falsification engine, not an optimizer.

Historical simulation is only trusted here to KILL ideas. A good backtest makes an idea eligible for live paper trading; it is never treated as a prediction of profit. The strongest force in quantitative research is self-deception, and this rule is the main defense.

Rules are written before results exist.

Pass/fail criteria — including the December go-live gate and each experiment's test battery — are written down and frozen before the data that will judge them arrives. Deciding the rules after seeing the outcome is how every fake track record in this industry gets made.

Tests run on survivorship-free data.

Backtests here include the stocks that later went bankrupt or were delisted — the ones most datasets quietly forget. Including the corpses typically makes results substantially worse, and substantially more honest.

Headline numbers ship with their asterisks.

Example, applied to our own flagship: the momentum backtest's headline is the LUCKIEST of eight calendar start-date variants — across all eight, the average result is roughly zero. We publish that spread alongside the headline, because a number without its luck estimate is advertising, not research.

A negative control runs live.

One sleeve trades daily under a strategy our tests rejected, with a covenant that it can never be promoted. If it quietly wins, our tests are broken. No result from the other sleeves is fully trusted unless the control behaves as expected.

the journal

Every research document, published as written

The live research journal: every internal research document, published verbatim (credential identifiers redacted, count shown) by the same automated daily export that publishes the trading record. Nothing here is published by hand; a document committed to the repository appears on the next daily cycle, and a tested completeness invariant guarantees none is silently missing.

DISMISSED — capture bid/ask/sizes/quote-timestamp at order-submit time as a new forward streamliving docDISMISSED — SV-1 — the state schema version is un-bumpable on the two rewrite-in-place filesliving docSTAMP-3 — #132's pre-open→previous-day rule inverts for current-state captures (shortability), and three stamp sites still judge the UTC dateliving docSTAMP-2 — the trading path shares STAMP-1's UTC-date defect: a midnight-crossing delayed run consumes the NEXT session's idempotency slot, the real session is silently skipped, and the no-run canary reads greenliving docSTAMP-1 — the capture writers use the runner's UTC date as BOTH the trading-day stamp and the idempotency keyliving docSKW-4 — the skew store serves 2,068 rows stamped on non-sessions (two NYSE holidays and a Saturday) through a reader with no trading-day filter, and per-name coverage is unmeasured: GHC and NVR have zero rows in 62 capture days and no instrument would say soliving docSKW-3 — the backup firing silently replaces in-session term-structure rows with post-close quotes stamped `healed=False`: the skew job's second-firing idempotency has a per-stream hole, and capture-QA cannot see itliving docSKW-2 — the skew capture sends dash-class symbols Alpaca rejects; BRK-B and BF-B have never capturedliving docSKW-1 — the fourth Alpaca client has no request deadline: the skew-capture cron can still hang foreverliving docSAC-1 — census of trading-path persist sites: "state advanced or keyed without regard to run completion" (member four already found)living docRWG-1 — the runaway-order guard: phantom turnover-blend base on a HALT, and a blocked drawdown liquidationliving docRPL-1 — the marketable-limit market-replace is the one order path outside the safety chokepointliving docPIT-1 — the pit snapshotter greens while producing nothing (or garbage that permanently occupies its slot): three silent shapes, bleeding under the displaced-cron regimeliving docPC-1 — ceiling on the cap-explained book (owner policy; now with a measured consequence)living docOSR-1 — 130 market replacements were REJECTED by the broker after the local log recorded them ACCEPTED, and no channel can see a broker-side terminal rejection: the order log is write-once at submit, the digest counts local `failed` only, reconcile compares ID setsliving docMEM-1 — the `memory/` directory does not exist in the repo, and the C0 kill list is one of the files that is missingliving docDISMISSED — LCK-1 — the CI lock is version-pinned but hash-lessliving docKSP-1 — the kill-switch day erases the sizer's memory of the crash: the crash holding period never reaches the Kelly ledger (the validated engine books it), and the risk-off flag every later vol estimate trusts is recorded intent, not the observed bookliving docKBM-1 — dismissed: a dateless Kelly snapshot would silently skip booking forever (unreachable via any current writer)living docIGD-1 — five guards that exist and cannot fire: a bool-returning call wrapped in an unreachable `except` with the return discarded — at the ONE submit site outside the safety gate, the mid-run halt, and three alert emails; the test that "pins" the first one teaches a mock a failure no broker has; a run whose orders all failed exits 0living docHIRE-1 — headcount as a new data dimension: does hiring growth (a documented anomaly) survive on this apparatus, and is revenue-per-employee worth a look?living docGAP-1 — six permanent capture-days are already lost, and charter gauge G4 claims they are machine-checked when no instrument can see themliving docFOS-1 — a per-symbol failed-order streak is indistinguishable from a one-off failure: AVB has failed 9 consecutive sessions, six of them the EMERGENCY-EXIT path, and no channel escalatesliving docDISMISSED — F2-REOPEN — the fleet email's queue-count parserliving docEXQ-1 — one invalid symbol dumps the ENTIRE order batch to market orders (live momentum exposed; contaminates the December TCA gate)living doccron-delay-capture — capturing GitHub-Actions cron-delay telemetry as a forward stream: considered and REJECTEDliving docCQA-2 — capture-qa reads only the latest file, so a missing today-file fails nothing and a failing stream mails under another stream's nameliving docCNT-1 — the go-live gate progress PUBLISHED on thales.report is one ahead of what the gate's own evaluator computesliving docDISMISSED — CLK-1 — two dated criteria mature with no evaluator, and neither is on the NORTHSTAR §6 calendarliving docDISMISSED — B4-NOTCH — the skew stream's calendar notch, and the ≥400-row floor it put in questionliving docDISMISSED — AUTO-APPROVE — pre-register which classes the machine may approve without the ownerliving docAMG-5 — the new sanctioned-move guard (47e8e73) fails OPEN three ways; one lets an unrecognized git file-status write into the owner's approval gateliving docAMG-4 — the records auto-merge workflow judges a PR by the PR's OWN copy of the guard, so its "can never widen itself" claim is circularliving docDISMISSED — AMG-3 — the fenced-real-gate / plain-decoy self-approval in the records auto-merge guardliving docRESEARCH_QUEUE.md — FROZEN ARCHIVE (2026-08-21)living docRESEARCH_AUDIT_LOG.md — the pipeline's own measurement seriesliving docThales Research Notebook3 redactedliving docThales Research Notebook1 redactedliving docqueue/approvedliving docqueue/builtliving docqueue/dismissedliving docqueue/openliving docNORTHSTAR.md — The Self-Improvement Charterliving docThales — design review against first principlesSep 5, 20262026-09-05 — External design review (r2): maintainer-side verificationSep 5, 20262026-09-02 — Deep self-audit (interactive, fresh-context)Sep 2, 20262026-08-31 — The displaced-cron regime: three days of trading after the closeAug 31, 2026Design review — thalesAug 29, 20262026-08-29 — External design review: maintainer-side verificationAug 29, 20262026-08-21 — meanrev clearing-day churn: the RWG-1 leg-2 measurementAug 21, 2026Routine self-sustainment audit — the loops did not close (2026-08-08)Aug 8, 20262026-08-07 — the skew capture lost a day and gained a contaminated oneAug 7, 20262026-08-06 — the day the fleet flew blind, and nobody was toldAug 6, 20262026-08-04 — the live momentum book has been 24% larger than its design, silentlyAug 4, 2026VRP v2 — pre-registered spec (FROZEN AT COMMIT, 2026-08-01)Aug 1, 2026Quant panel - suggestions report (2026-08-01)Aug 1, 2026VRP pessimistic twin — same-snapshot fix (panel A1) — 2026-08-01Aug 1, 2026External design review (2026-08-01) — archived with ground-truth correctionsAug 1, 2026December governance package — DRAFT (2026-08-01)Aug 1, 2026Bounding the Stateful-Purge Leak — Pre-Registration (2026-08-01)Aug 1, 2026Null Calibration of the Falsification Engine — Pre-Registration (2026-08-01)Aug 1, 2026C1 — Go-live leverage diff: cap gross exposure, not the scalarAug 1, 2026Turn-of-Month Forward Attribution — Pin (pre-registered 2026-08-01, panel item B5)Aug 1, 2026Skew Stream — Pinned Activation Test (pre-registered 2026-08-01, panel item B4)Aug 1, 2026Adjusted-opens audit — 2026-08-01Aug 1, 2026A4 — Publication-lagged days-to-cover exclusion battery (pre-registration)Aug 1, 2026Short-flow → short-interest nowcast — ONE-SHOT FALSIFICATION (panel A3, 2026-08-01)Aug 1, 2026VRP decision-fidelity amendment — 2026-07-25Jul 25, 2026Micro-cap momentum battery — PRE-REGISTRATION v1 (2026-07-22)Jul 22, 2026Micro/small-cap momentum — data feasibility scoping (2026-07-21)Jul 21, 2026December Gate — Interpretation Memo (pre-registered 2026-07-13)Jul 13, 2026Multi-Sleeve Platform — SPEC (pre-registered 2026-07-05)Jul 5, 2026OOS Degradation Monitor — Power & Detection-Latency Study (2026-05-30)May 30, 2026Capacity & Stress-Execution Realism — 2026-05-30May 30, 2026Apparatus Readiness — 2026-05-30 AFK session indexMay 30, 2026Post-Refresh Fundamentals-Signal Test ProtocolMay 29, 2026Execution-Fidelity Fix Spec — unify backtest & live constructionMay 29, 20262026-04-06 Research NotesApr 6, 2026

These are internal working documents — experiment pre-registrations, incident write-ups, verdicts including nulls — shipped verbatim by the same automated daily export that publishes the trading record. Machine-readable index: /data/research.json.

the kill list

Built, tested, dead

Every strategy idea this platform built and tested, with the honest verdict. Publishing failures is the point: a research process that only shows winners is indistinguishable from luck. Verdicts come from pre-registered statistical gates (overfitting probability, out-of-sample tests on survivorship-free data), not from taste. Much of this list was built, tested, and killed by the AI agent that operates the platform, during overnight research sessions run under the same pre-registered rules.

Machine-learning trade filter (random forest meta-labeling)2026-05Killed

Overfitting probability measured at 100% with price-only features — the model memorized the past instead of learning anything. Disabled in config; would only be revisited with fundamentally new input data.

Residual momentum (factor-neutralized ranking)2026-05Killed

Worked mechanically exactly as the literature describes — and cost about 1.7% of annual return in practice. An idea can be real and still not worth its price.

Value sleeve (cheap-stock tilt)2026-05Killed

Added no robustness benefit on the overfitting tests. Removed.

Quality, accruals, asset-growth and low-volatility tilts2026-05Killed

Every classic fundamental tilt was built and tested; none survived the honest evaluation harness. The conclusion — published price+fundamental signals on US large caps are exhausted — is now a standing research rule here.

Graduated index-hedge overlay (scaling into SPY in stress)2026-05Killed

A return drag: about −1% annualized in walk-forward testing, losing in 0 of 5 test windows. Protection that costs more than the crashes it softens.

A faster variant of the flagship's signal2026-05Not shipped

Worsened the overfitting probability in both test modes. The incumbent stayed.

A three-parameter variant of the flagship2026-05Not shipped

Looked better on the headline number, but the verdict flipped depending on how the test data was partitioned — the signature of a fragile result. Investigated, documented, not shipped.

Tranching (staggering the rebalance across the month)Jun 10, 2026Killed

Killed by its own pre-registered gate before launch. Its useful by-product survived: a diagnostic that measures how much of any backtest headline is start-date luck — now run against our own numbers.

Corporate-insider trading signal (SEC Form 4 filings)Jun 10, 2026Killed

A perfect null: after building the full data pipeline, insider filings predicted nothing at all in our universe. The cleanest kind of kill.

Short-horizon mean reversionJul 5, 2026Repurposed

Rejected by its pre-registered test battery — then deliberately deployed live anyway as the negative control, under a covenant that it can never be promoted. If our rejected strategy quietly wins, our tests are broken. It trades every day to keep that check honest.

Micro-cap reversal program (v1)Jul 22, 2026Killed

A month of data engineering (point-in-time small-cap universe, delisted-stock recovery, realistic cost model), then one hash-frozen test battery allowed to run exactly once. It failed: out-of-sample Sharpe −0.30, negative even at half the estimated trading costs. Killed without appeal; any v2 requires a brand-new pre-registration.

Learning-to-rank stock selection (LambdaRank)2026-05Paused

No honest gain over the simple momentum ranking to date. Archived, not deleted — would be revisited only with genuinely new input data.

Congressional-trading data (as a signal source)Jul 25, 2026Killed

Evaluated during a data-sourcing deep-research pass: since the 2012 disclosure law, politicians' disclosed stock picks slightly underperform. A famous signal that died the moment it became famous.

still standing

What survived

The configuration running in production today.

What survived every gate and runs in production today is a single trend-following configuration on large-cap US equities, wrapped in layered risk controls: diversification-aware weighting, conservative sizing, volatility targeting, and a hard drawdown stop. The exact signal and parameters stay in the private repository — what we publish is every output it produces, every day, and the honest scorecard including its caveats.