# Hypothesis lab — what to test next, and how (deep research, 2026-09-28)

> Owner request (2026-09-27): *"be a little more intentional about what kinds of experiments
> to run and what combinations to produce … do some deep research … pull in some more
> expertise."* Produced by the `deep-research` workflow (fan-out search → fetch → claims
> extracted → 3-vote adversarial verification → synthesis): 24 sources fetched, 118 claims extracted, 25 put to verification, 18 confirmed, 7 refuted; 106 agent calls.
> Findings, evidence and caveats below are the workflow's verified output, reproduced verbatim;
> the power arithmetic in finding 7 was independently re-checked by simulation (formula t vs
> simulated t: 0.67/0.64, 1.41/1.38, 3.11/3.07). The final section is a PROPOSAL, not a decision.

## Summary

The verified research explains the lab's null results; it does not contradict them. The timing ideas proposed most often have all failed: macro and valuation signals (credit and term spreads, dividend and earnings yields, CAPE, including combined and theory-restricted versions) and calendar rules. They faded once their original samples were extended, failed after correcting for how many had been tried, or never beat buy-and-hold, and most of their historical success came before the mid-1970s to early 1990s. Only two families have durable long-run support, and neither maps cleanly onto a long-only switch between ETFs. One is trend following run as a diversified long/short book across 67 futures markets: a net Sharpe ratio (return per unit of risk) of 0.76 over 1880-2016, but 0.41 in 2010-16. The other is volatility targeting of stocks and credit: +0.08 to +0.11 Sharpe over about 90 years, nothing for bonds or commodities, not significant in direct tests, and roughly zero after 1936. Realistic timing gains of about +0.1 Sharpe cannot be seen in an 18-year window; a 90-year test of one gave p = 0.30. By this synthesis's own rough calculation, the lab's current lab-wide bar (about 3.5-4 standard errors after 1,344 trials) can only certify gains of roughly +0.5 Sharpe or more on 2005-2022, so the nulls were close to guaranteed whether or not modest real effects exist. The next studies should therefore buy statistical power instead of producing more variants. That means a handful of pre-registered rules with parameters taken from the literature, tested first on long pre-ETF histories (Ken French daily data from 1926, bond returns rebuilt from Treasury yields) and in other countries. Each should be judged on consistency across sub-periods and on flat parameter plateaus, under an error budget split in advance between confirmatory and exploratory work.

## Findings

### 1. (high confidence)

Timing the stock market with macro and valuation signals has the weakest record of anything reviewed. The family covers dividend and earnings yields, CAPE, the term spread, the credit (default) spread and related series, and is closely related to the lab's credit, yield-curve and jobless-claims dashboards. Published predictors faded when their samples were extended, and almost all fail a data-snooping correction. The standard rescues stopped working after about 1993: combining forecasts, theory-based sign and zero-floor restrictions, penalized regression and principal components. Timing strategies built on these signals, including CAPE timing over 116 years, did not beat buy-and-hold.

<details><summary>Evidence</summary>

DECAY (Goyal-Welch-Zafirov, RFS 2024): 29 predictors from 26 papers published after 2008 were re-run to Dec 2021, typically about 10 extra years that still include the discovery data. 12 of 29 lost even in-sample significance. Half of the rest did poorly out of sample (on data not used to fit them). 69% saw their t-statistics fall. The authors' simulations fit 'all spurious from the start' better than 'a modest stable effect' (their wording is hedged).

COMBINING DOES NOT RESCUE IT (same paper): a t-weighted consensus forecast had an out-of-sample R² (gain over simply forecasting the historical average) of 3.68% on annual data only if data were assumed to be released instantly. With a realistic 6-month release lag it was -0.73%. All of its good performance came before the mid-1970s, it scored -0.8% in 2012-2021, and every strategy built on it did worse than holding stocks all the time.

REPLICATION FAILURE (Denk-Löffler, RAPS 2024; figures from the 2022 working paper):
- The Rapach-Strauss-Zhou mean combination had R² 5.07% and a +3.49%/yr utility gain (net of 20 bp) in 1965-92, then -0.38% and -0.28%/yr in 1993-2020.
- Every Rapach-Strauss-Zhou scheme was negative on both measures after 1992.
- No method predicts in recent decades: combinations, ridge/LASSO/elastic net, principal components, three-pass filter.
- With Campbell-Thompson restrictions nothing is significant at 5%; the best is R² 1.20% with a +0.01%/yr gain.
- The authors' named cause: the errors of individual forecasts became more correlated, so combining them stopped diversifying.
- The break predates the 2010 publication by about 17 years, so this is instability over time rather than classic post-publication decay.

CORRECTION FOR THE SEARCH (Dichtl et al., IJF 2021): each strategy tested had been published as beating the historical mean. Almost all fail out of sample once the whole set is corrected for data snooping. The abstract excepts 'only few' sum-of-the-parts variants; that exception is contested (see caveats).

STRATEGY LEVEL (Asness-Ilmanen-Maloney, JoIM 2017): real-time CAPE timing with 50-150% equity over 1900-2015 had a Sharpe of 0.37 vs 0.38 for buy-and-hold (0.37 vs 0.37 in 1958-2015). That is before costs, and neither difference is significant. Because valuations drifted up, the rule averaged only 89% invested after 1958, about -0.6%/yr of forgone equity premium. Trading only at extremes did not meaningfully help.

LAB: consistent with the NULL on the 'macro cycle' and 'fear + credit' dashboards in comprehensive_v1.

Sources: https://academic.oup.com/rfs/article/37/11/3490/7749383, https://academic.oup.com/raps/article-abstract/14/4/545/7643730, https://www.sciencedirect.com/science/article/abs/pii/S0169207020300510, https://www.aqr.com/library/journal-articles/market-timing-sin-a-little

</details>

### 2. (high confidence)

Volatility targeting (cutting exposure when recent volatility is high) gives stocks and credit a small Sharpe gain over about 90 years, flat across settings. It does nothing or harms bonds, currencies and commodities. Even for US stocks the gain is statistically marginal in direct comparisons, concentrated around the Great Depression, and roughly zero after 1936 and over 2005-2022.

<details><summary>Evidence</summary>

HARVEY ET AL. (JPM 2018, Man AHL authors), US stocks 1927-2017 at a 10% target:
- Sharpe went from 0.40 to 0.48-0.51.
- The result is flat across EWMA (exponentially weighted moving average) half-lives of 10-90 days, and gross is about equal to net at ~1 bp costs.
- Alpha of scaled over unscaled: 0.64 bp/day, t = 3.05, R² 0.73.
- It failed in the 1958-87 third of the sample.
- S&P 500 futures 1988-2017: 0.50 to 0.57-0.60 (gross, using daily-data volatility).
- A 60/40 portfolio went from 0.80 to 0.87-0.91 (verifier note).

OUTSIDE RISK ASSETS (same paper): US 10-year bonds 1963-2017 (returns rebuilt from yields) fell from 0.25 to 0.05-0.09, all because of the pre-1980 bond bear market. 10-year Treasury futures: 0.64 vs 0.63. Across 50 futures and forwards, only equity indices and credit improved, and only slightly. The mechanism: next-month stock returns show no clear pattern across volatility quintiles, but bond returns were much higher in the most volatile quintile. Cutting exposure when volatility is high therefore throws away bonds' best months.

CEDERBURG-O'DOHERTY-WANG-YAN (JFE 2020): 103 US equity strategies, mostly long-short decile spreads, with Moreira-Muir monthly variance scaling.
- 53 improve and 50 get worse (binomial p 0.84); 8 are significant wins and 4 significant losses.
- So the result does not replicate outside momentum, where all 9 variants improve and 5 significantly.
- The in-sample 'spanning' alphas do replicate (77 of 103 positive, 23 significant), but capturing them needs hindsight weights.
- A real-time version (expanding window starting with 10 years, risk aversion 5, leverage up to 5x) earned a lower certainty-equivalent return (the safe return an investor would accept instead) in 72 of 103 cases. Only 45 of 103 had a higher Sharpe. The result holds across windows, risk aversion and leverage caps from 1x to unlimited, and is before costs.
- The US market itself, 1926-2016: 0.51 vs 0.42, p = 0.30. The real-time version: 0.42 vs 0.46. The gain is concentrated around the Depression.

VERIFIER RERUNS ON KEN FRENCH DATA:
- After Aug 1936: 0.49 vs 0.49.
- 2005-2022: 0.55 vs 0.55.
- A daily, unlevered version (capped at 100% invested; 12% target; 20-day half-life) on 2005-2022: +0.10 Sharpe, p about 0.34-0.40.

ASIDE FOR THE MOMENTUM SLEEVE: momentum portfolios were the one place volatility management reliably helped. These were long-short factor portfolios, so this is only suggestive for Thales' long-only book.

Sources: https://people.duke.edu/~charvey/Research/Published_Papers/P135_The_impact_of.pdf, https://www.lehigh.edu/~xuy219/research/COWY.pdf

</details>

### 3. (medium confidence)

Diversified time-series trend following has the longest out-of-sample record of any family reviewed. But the record is for a long/short, volatility-scaled book across 67 futures markets, not a long-only switch on one index. It weakened sharply in 2010-2016, most of all in its fast (1- and 3-month) signals.

<details><summary>Evidence</summary>

THE RECORD (Hurst, Ooi & Pedersen, JPM 2017, AQR): Jan 1880 to Dec 2016, an equal blend of 1-, 3- and 12-month trend signals across 29 commodities, 11 equity indices, 15 bond markets and 12 currency pairs, scaled to 10% volatility.
- 7.3%/yr after simulated costs and 2/20 fees (11.0% after costs only; 18.0% before both), at 9.7% volatility: a net Sharpe of 0.76.
- Correlation -0.01 to US stocks and -0.03 to US 10-year bonds.
- Positive in every decade.
- The pre-1985 century is presented as out of sample relative to Moskowitz-Ooi-Pedersen (2012). Strictly, only about 1880-1964 is untouched, because MOP used data from 1965.

2010-2016 WEAKNESS:
- Net Sharpe 0.41 (3.3%/yr at 8.1% volatility), the weakest since the 1910s.
- Gross Sharpe by signal fell from 1.38 to 0.06 (1-month), from 1.19 to 0.30 (3-month) and from 1.32 to 0.73 (12-month); even the 12-month figure is that signal's worst period.
- The authors link the weakness to high cross-asset correlation (markets moving together as 'risk-on/risk-off', late 2008 to mid 2014), not to publication.
- AQR's 2020 follow-up points instead to unusually muted market moves (verifier note).

CAVEATS:
- Before futures existed, returns are cash indices financed at local short rates, and the authors say the strategy was not implementable early on.
- The net figures depend on assumed costs, set at 6x modern estimates before 1993.
- The authors are AQR principals, and AQR sells managed futures.

CRITIQUES SEEN ONLY IN VERIFIER NOTES:
- Huang-Li-Wang-Zhou (JFE 2020, 'Time-series momentum: is it there?') find weak asset-by-asset predictability while conceding the strategy is profitable. Much of the profit can come from simply being long assets with high average returns, which is exactly what the lab's Q2 test (timing vs a static mix) strips out.
- Kim-Tse-Wald (2016) credit the alpha to volatility scaling.

LAB: the Sharpe here comes from breadth (many weakly correlated trends) and from the short side. The lab's single-index 200-day SPY rule matched SPY's Sharpe (0.47-0.50 vs 0.47), consistent with capturing little of either.

Sources: https://www.aqr.com/Insights/Research/Journal-Article/A-Century-of-Evidence-on-Trend-Following-Investing

</details>

### 4. (medium confidence)

Carry means holding what pays more to hold and shorting what pays less. It is a strong cross-asset factor in-sample, but the evidence is long/short, before costs and ends in 2012, with no later test period. A long-only ETF version keeps only the long leg, and the claim that carry works as a single-asset timing signal failed verification.

<details><summary>Evidence</summary>

KOIJEN, MOSKOWITZ, PEDERSEN & VRUGT (NBER w19325, Aug 2013; the JFE 2018 version was not checked):
- Average carry Sharpe within an asset class is about 0.74. Excluding index-put carry, whose 1.80 lifts the average, it is about 0.64.
- A diversified 'global carry factor' scores 1.10 vs 0.47 for a passive long portfolio weighted the same way.
- Carry vs passive by class: equity indices 0.88 vs 0.32, commodities 0.60 vs 0.08, currencies 0.68 vs 0.36, US Treasuries across maturities 0.68 vs 0.57, credit 0.47 vs 0.34.
- Carry LOST to passive in global 10-year bonds (0.52 vs 0.74) and 10y-2y slopes (0.66 vs 0.71).

LIMITS:
- Samples start in 1971-1996 and end in 2011-2012.
- The paper's 'out-of-sample' tests are new asset classes, not new years.
- Cross-class weights use full-sample volatility, which is look-ahead.
- No after-cost returns are reported.
- Every strategy is a zero-cost long/short.
- The headline moved across versions (1.49 in a 2012 draft).

The companion claim that carry works as a long-or-short timing signal on each asset (Sharpe 0.40-0.78 by class) was refuted 0-3 and is not used here.

LAB: the nearest long-only version with a long sample is choosing Treasury maturity by carry. It can be tested from 1962 with returns rebuilt from FRED yields.

Sources: https://www.nber.org/system/files/working_papers/w19325/w19325.pdf

</details>

### 5. (high confidence)

Searching through calendar rules is a textbook data-snooping trap. The best of 9,452 calendar rules on a century of Dow data was not significant once the size of the search was accounted for. The famous Monday effect faded after publication, lost money out of sample and could never have covered trading costs.

<details><summary>Evidence</summary>

THE SEARCH (Sullivan, Timmermann & White, UCSD working paper 1998; J. Econometrics 2001): 9,452 rules on the DJIA, 1897-1996, covering day-of-week, week-of-month, month-of-year, semi-month, holiday, end-of-December and turn-of-month.
- The best rule was 'out of the market on Mondays': 8.66%/yr vs 4.63%, nominal p < 0.002.
- White's Reality Check p, which accounts for the whole search, was 0.24 (0.20 for 1897-1986).
- Under a Sharpe criterion the Reality Check p was 0.47-0.55.

AFTER PUBLICATION:
- From its publication (Cross 1973) to 1996: 9.02% vs 8.78%, nominal p 0.44.
- From June 1986 to 1996 it lost: 10.2% vs 11.6%/yr.

VERIFIER'S INDEPENDENT CHECK (CRSP daily returns):
- Monday's mean return was -17.6 bp/day (t -7.4) in 1926-72 but -2.7 bp (t -0.87) in 1973-96.
- The rule underperformed in 1987-96 and again in 1997-2024.
- Zakamulin (2026) finds no Monday effect in 1994-2024.
- At about 104 round trips a year, 5 bp one-way costs come to about 10%/yr.

DILUTION: the same paper shows that padding a family weakens the test. Adding semi-month rules raised the best rule's Reality Check p from about 0.33 to about 0.52 without changing the rule.

SCOPE: no claims about pre-FOMC drift, turn-of-month or Halloween survived this review. Hansen-Lunde-Nason (2005, verifier note) found some end-of-year effects significant across 10 national indices with a more powerful test, weaker since the late 1980s.

Sources: https://escholarship.org/uc/item/2z02z6d9

</details>

### 6. (high confidence)

Realistic timing gains (roughly +0.05 to +0.10 Sharpe) cannot be reliably detected in an ~18-year window. A null result there is weak evidence against a small real effect. About 90 years of daily data only reached t ≈ 3 for volatility targeting, and the direct Sharpe comparison over 1926-2016 gave p = 0.30. A predictor that is genuinely true often loses to the historical average over an entire test window.

<details><summary>Evidence</summary>

HARVEY ET AL.: alpha t = 3.05 over about 91 years of daily data. Scaled down to 18 years (multiply by √(18/91)) that gives t ≈ 1.36, about 27% power at 5% two-sided. This is the verifier's calculation and assumes a constant effect.

CEDERBURG ET AL.: +0.09 Sharpe for the scaled market, p = 0.30 over 90 years. Verifier reruns on 2005-2022 gave p of about 0.34-0.99, depending on the version.

CAMPBELL & THOMPSON (NBER WP 11468, 2005, Table 4; 5,000 simulations calibrated to each predictor's full-sample fit):
- Even with the true coefficient known, the predictor loses to the historical mean on forecast error in 4.9% (smoothed earnings/price) to 38.4% (payout ratio) of samples. Examples: term spread 23.8%, default spread 28.0%, book-to-market 9.5%.
- It gives the investor lower utility in 6.1-54.4% of samples.
- With estimated coefficients and both restrictions, it loses in 13.4-62.5%.
- This simulation was dropped from the published 2008 RFS version.
- The simulated model favours the predictor, so true loss rates are likely higher.
- For book-to-market and the default spread, the paper judged the observed failures bigger than bad luck explains.

GOYAL-WELCH-ZAFIROV: merely picking the best of 3 data frequencies lifts the honest 10% critical t from 1.65 to ≈ 2.1.

IMPLICATION: the lab's nulls close the door on LARGE effects only. They cannot rule out effects of the size the literature documents.

Sources: https://people.duke.edu/~charvey/Research/Published_Papers/P135_The_impact_of.pdf, https://www.lehigh.edu/~xuy219/research/COWY.pdf, https://www.nber.org/papers/w11468, https://academic.oup.com/rfs/article-abstract/21/4/1509/1567518, https://academic.oup.com/rfs/article/37/11/3490/7749383

</details>

### 7. (medium confidence)

By a standard rough calculation, the lab's 2005-2022 window under its current lab-wide bar can only certify timing gains of about +0.5 Sharpe or more, five times the best-documented effects. Even a century of data certifies only about +0.2. A +0.1 gain needs about 250 years of data for an 80% chance of detection even at a plain 5% level.

<details><summary>Evidence</summary>

This is the synthesizer's calculation, not a verified claim.

THE FORMULA: the lab's Q2 test compares a rule with its matched static mix at equal volatility. Its t-statistic is roughly ΔSR × √years ÷ √(2(1-ρ)).
- ΔSR is the annual Sharpe difference.
- ρ is the daily correlation between the rule and its mix, about 0.75-0.85 for a switch that is invested about 70% of the time. For example, a rule that holds stocks 70% of the time and cash otherwise has ρ ≈ √0.7 ≈ 0.84.
- It assumes roughly normal, serially independent differences. Fat tails and autocorrelation, which the lab's block bootstrap handles, make things somewhat worse.

AT ρ = 0.8:
- A +0.10 gain gives t ≈ 0.67 in 18 years, t ≈ 1.4 in 79 years (1926-2004) and t ≈ 1.6 in 100 years.
- Years needed for 80% power: for +0.10, about 250 at 5% one-sided, about 400 at 1% and about 730-920 at the lab-wide bar. For +0.20: about 60 / 100 / 180-230. For +0.30: about 27 / 45 / 80-100.
- In the lab's 18-year window, a +0.1 gain is detected 16% of the time at a plain 5% and essentially never at the lab-wide bar. Even +0.3 is detected only 3-8% of the time at the lab-wide bar.
- The smallest gain detected half the time: over 18 years, +0.25 at plain 5% and +0.51-0.59 lab-wide. Over about 100 years, +0.10 at plain 5% and +0.22-0.25 lab-wide.

THE LAB-WIDE BAR: study.py multiplies each study's Romano-Wolf p (the lab's within-study correction) by (all lab trials) ÷ (this study's rules). With 1,344 trials charged, a new one-rule study must reach p ≤ 0.05/1,345 ≈ 3.7e-5, which is about 3.96 standard errors one-sided. A large study of correlated rules lands nearer 3.4.

SANITY CHECKS:
- The formula roughly reproduces published p-values: Cederburg et al.'s +0.09 over 90 years comes out at p ≈ 0.2-0.3.
- It also matches the lab's own result. comprehensive_v1's best allocator was +0.17 over its matched mix across about 17 years, a single-rule p ≈ 0.13. As the best of 408 rules, its timing p was 0.92.

Sources: /Users/dan/Documents/GitHub/thales/src/thales/lab/study.py, /Users/dan/Documents/GitHub/thales/research/lab/README.md, https://www.lehigh.edu/~xuy219/research/COWY.pdf, https://people.duke.edu/~charvey/Research/Published_Papers/P135_The_impact_of.pdf

</details>

### 8. (medium confidence)

What genuinely raises power is more independent data per hypothesis and fewer hypotheses sharing the error budget. That means longer pre-ETF histories, returns rebuilt from yields, other countries as replications, and effects that recur many times a year. Theory-based restrictions shrink the number of trials but did not rescue weak predictors.

<details><summary>Evidence</summary>

PRECEDENT (verified): the literature tests exactly these rules on histories the lab does not hold:
- Ken French daily US market data from 1926 (Harvey et al., Cederburg et al., the verifier's reruns).
- Shiller's monthly data (Asness et al., 1900-2015).
- Cash indices from 1880 across 67 markets (Hurst et al.).
- A century of daily Dow data (Sullivan et al.).
- Bond returns rebuilt from Treasury yields (Harvey et al., 1963-2017).

WHAT THE LAB HOLDS (repo):
- data/factors/ff5_daily.parquet starts 1963-07. The five-factor file has no earlier data; Ken French's three-factor daily file starts 1926, and his industry portfolios could do the same for sector rotation. momentum_daily starts 1926-11.
- data/lab/cash_DTB3 starts 1990.
- The lab's signal snapshots (VIX, 10y-2y, Baa-Aaa, claims) start 2000. Most of these series exist years to decades earlier at FRED; VIX starts only in 1990.
- Studies are scored on one common window set by the latest-starting ETF (comprehensive_v1 began 2005-11). Yet six country ETFs date from 1996, three emerging-market ETFs from 2000 and the sector SPDRs from 1998. Single-market replications run on each fund's own window would add years.

OTHER COUNTRIES (verifier notes): Bongaerts-Kang-van Dijk (FAJ 2020) found an implementable US vol-targeting gain of +0.15 Sharpe (1982-2019) that averaged only +0.04 across 10 markets. Replication abroad is a cheap way to catch US-only flukes, though 10 correlated markets are fewer than 10 independent tests.

NUMBER OF BETS (inference): a rule that bets about 50 times a year, like the Monday rule, piles up independent evidence far faster than a 12-month trend switch. The trend switch changes state only a few times a year, so it makes a few dozen independent bets over 18 years whatever the data frequency.

RESTRICTIONS: sign and zero-floor restrictions did not make macro forecasts significant after 1993 (Denk-Löffler). The claim that restrictions rescue many predictors failed verification (1-2). Their value is fewer free parameters, not a stronger effect.

Sources: https://people.duke.edu/~charvey/Research/Published_Papers/P135_The_impact_of.pdf, https://www.lehigh.edu/~xuy219/research/COWY.pdf, https://www.aqr.com/Insights/Research/Journal-Article/A-Century-of-Evidence-on-Trend-Following-Investing, https://www.aqr.com/library/journal-articles/market-timing-sin-a-little, https://escholarship.org/uc/item/2z02z6d9, https://academic.oup.com/raps/article-abstract/14/4/545/7643730, /Users/dan/Documents/GitHub/thales/data/factors/ff5_daily.parquet, /Users/dan/Documents/GitHub/thales/data/lab/signals, /Users/dan/Documents/GitHub/thales/config/lab.yaml

</details>

### 9. (medium confidence)

The verified studies point to a concrete design discipline:
- take parameters from prior work and demand a plateau around them, not a grid optimum;
- price every hidden choice;
- keep families tight;
- require the same sign across sub-periods and survival without the single best episode;
- combine only signals whose errors are weakly correlated;
- check a mechanism's precondition before testing the rule.

<details><summary>Evidence</summary>

PLATEAUS: the vol-targeting gain is within 0.03 Sharpe (0.48-0.51) across 10-90-day half-lives, and gross is about equal to net (Harvey et al.). AQR's trend book uses a fixed equal blend of 1/3/12-month signals, not a fitted lookback (Hurst et al.).

HIDDEN CHOICES COST SIGNIFICANCE:
- Choosing the best of 3 data frequencies moves the 10% critical t from 1.65 to ≈ 2.1 (Goyal-Welch-Zafirov).
- Real-time details flip results. The consensus forecast's R² goes from +3.68% to -0.73% with a 6-month data-release lag. Cederburg et al.'s real-time vol-managed version loses where the hindsight version wins.
- The lab's lag_days convention is the right defence. Latest-vintage revisions to jobless claims remain a small leak, per the lab README.

TIGHT FAMILIES: in Sullivan et al., adding semi-month rules raised the best rule's Reality Check p from about 0.33 to about 0.52 without changing the rule (verifier's reading). Padding a family with low-prior variants spends power. Hansen's SPA test, which the lab uses, reduces this effect but does not remove it.

SUB-PERIODS AND EPISODES: every verified effect was unstable or concentrated in one stretch.
- Vol targeting failed in 1958-87, and its market gain sits around the Depression.
- Macro combinations worked only to the mid-1970s or early 1990s, much of it in 1973-75.
- Trend was weak in 2010-16.
- The Monday effect died after the mid-1980s.

COMBINATION: Denk-Löffler trace the collapse of combination forecasts to rising correlation among the individual forecasts' errors. For the lab (inference): dashboards whose trend, VIX and credit votes flip together in the same crises offer little diversification.

MECHANISM FIRST: vol scaling helps only where next-period returns do not rise with volatility (stocks) and hurts where they do (bonds). That precondition can be checked on long data before spending timing trials.

THRESHOLDS: stricter original t-hurdles only modestly predicted which predictors held up in Goyal-Welch-Zafirov: 36% improved at |t| ≥ 2.0, 43% at |t| ≥ 3.5, on only 7 variables. A stricter bar is no substitute for fresh data.

NOT COVERED BY VERIFIED CLAIMS: Harvey-Liu-Zhu's t > 3 guidance beyond this, specifics from López de Prado and from Bailey-Borwein-López de Prado-Zhu, alpha-investing, and holdout reuse.

Sources: https://people.duke.edu/~charvey/Research/Published_Papers/P135_The_impact_of.pdf, https://www.aqr.com/Insights/Research/Journal-Article/A-Century-of-Evidence-on-Trend-Following-Investing, https://academic.oup.com/rfs/article/37/11/3490/7749383, https://escholarship.org/uc/item/2z02z6d9, https://academic.oup.com/raps/article-abstract/14/4/545/7643730, https://www.lehigh.edu/~xuy219/research/COWY.pdf

</details>

### 10. (low confidence)

The lab's error budget currently gives every trial the same slice of 5% across the whole ledger. Wide exploratory grids therefore use up most of it. A single, well-motivated replication must clear about 4 standard errors, which is out of reach for realistic effects with any data that exists. Being 'intentional' can be written into the statistics by declaring in advance how the 5% is split between a few confirmatory tests and open-ended exploration.

<details><summary>Evidence</summary>

FROM THE REPO (not a verified claim):
- lab-wide p = min(1, Romano-Wolf p × lab trials ÷ study rules), in study.py, family_tests.
- This is a union bound, meaning it adds up each study's chance of a false alarm. It is equivalent to giving each study α × (its rules ÷ all trials): the same alpha per trial whatever the prior plausibility.
- Old runs are re-scored as the ledger grows.
- The vault separately spends 0.05/2^k on its k-th opening.
- With 1,344 trials charged, a one-rule study needs p ≤ 3.7e-5.

ILLUSTRATION (synthesizer's arithmetic): suppose 2.5% were reserved for at most five pre-declared confirmatory hypotheses (0.5% each, one-sided z ≈ 2.58) instead of 3.7e-5 each.
- That roughly halves the data needed for 80% power: (2.58+0.84)² ≈ 11.7 vs (3.96+0.84)² ≈ 23.1.
- A +0.2 Sharpe gain at ρ ≈ 0.8 over 1926-2022 would be detected about 70% of the time, versus about 20-37% under the current bar.
- Exploratory grids keep the other 2.5% under the existing machinery.

STANDARD ALTERNATIVES THIS REVIEW DID NOT VERIFY:
- weighted Bonferroni or alpha-spending with a pre-declared split;
- gatekeeping (test the family first, then its members);
- alpha-investing and online false-discovery-rate schemes that earn back budget on discoveries.

Any reallocation must be fixed and committed to the ledger before results are seen, or it becomes a forking path of its own.

CAVEAT: pre-1990 data is new to the lab but not to the literature, whose rules were partly designed on it. A confirmatory test there is a replication with more data, not a clean out-of-sample test.

Sources: /Users/dan/Documents/GitHub/thales/src/thales/lab/study.py, /Users/dan/Documents/GitHub/thales/research/lab/README.md

</details>

### 11. (high confidence)

Stop spending trials on:
- macro and valuation timing of stocks (credit and term spreads, yields, CAPE, claims dashboards);
- broad calendar grids;
- volatility scaling of bonds, gold or commodities;
- further rotation, canary and trend variants scored only on 2005-2022.

<details><summary>Evidence</summary>

Each kill condition is already met in the verified record.

MACRO AND VALUATION: no predictive ability in 1994-2022 by any method, restricted or not (Denk-Löffler). Published predictors decayed (Goyal-Welch-Zafirov). A data-snooping correction removes almost all of them (Dichtl et al.). CAPE timing scored 0.37 vs 0.38 over 116 years (Asness et al.). The lab's own dashboards came back NULL. At most one pre-registered sum-of-the-parts test is defensible, as Dichtl et al.'s contested exception.

CALENDAR GRIDS: the best of 9,452 rules had a Reality Check p of 0.24, and the Monday effect died after publication (Sullivan et al.).

VOL SCALING OF NON-RISK ASSETS: bonds fell from 0.25 to 0.05-0.09, and Treasury futures were unchanged (Harvey et al.).

NEW VARIANTS ON 2005-2022: realistic gains (about +0.1) are detected 16% of the time at a plain 5% in 18 years and essentially never at the lab-wide bar (see the power findings). Meanwhile every trial raises the bar for all future studies, so these variants buy almost no information at a real price.

Sources: https://academic.oup.com/rfs/article/37/11/3490/7749383, https://academic.oup.com/raps/article-abstract/14/4/545/7643730, https://www.sciencedirect.com/science/article/abs/pii/S0169207020300510, https://www.aqr.com/library/journal-articles/market-timing-sin-a-little, https://escholarship.org/uc/item/2z02z6d9, https://people.duke.edu/~charvey/Research/Published_Papers/P135_The_impact_of.pdf

</details>

### 12. (medium confidence)

Run next, in this order. Each item is one literature-fixed 'anchor' rule plus at most a few neighbours that move one setting at a time. The anchor is the test; the neighbours only check for a plateau.
1. A long-history replication of the lab's single-market trend rules on 1926-2004 US data.
2. One pre-specified, diversified long-only trend portfolio across asset classes.
3. Volatility management of the equity and credit sleeve only, judged as risk control.
4. Treasury maturity chosen by carry, from 1962.
5. An international replication gate for anything that passes.
6. At most one or two event-style calendar tests (turn-of-month, pre-FOMC), which this review could not support.

<details><summary>Evidence</summary>

1. SINGLE-MARKET TREND, LONG HISTORY.
- Rules: a 10-month moving average, 12-month time-series momentum (trailing excess return > 0) and the 1/3/12-month blend, each switching the US market into T-bills.
- Data: Ken French daily market and risk-free returns, 1926-2004. The lab has never searched this period, but the literature has used it, so this is a replication with more data, not a virgin test.
- Support: Hurst et al. The 12-month signal was the sturdiest in 2010-16.
- Expected gain over the matched mix: about 0 to +0.1. There is no verified single-market estimate: a claim that 12-month momentum timing lifted Sharpe from 0.38 to 0.48 over 1900-2015 was refuted 0-3. The lab's own 2005-2022 result was about 0.
- Pass: all of the following. Positive over 1926-2004. The same sign in at least 2 of 3 roughly 26-year thirds. Neighbours within about 0.03 Sharpe. Survives dropping its best episode (e.g. 1929-32). p under a pre-declared confirmatory bar.
- Kill: 0 or below over 1926-2004, or the wrong sign in 2 of 3 thirds. Then close the family, including any new 2005-2022 variants.
- Why first: it is cheap (3 rules) and adds 4-5x the years, so it settles whether the ETF nulls came from low power or from no effect.

2. DIVERSIFIED LONG-ONLY TREND (the long-only cousin of Hurst et al.).
- Design: 8-10 sleeves (US, developed ex-US and emerging-market equity; 7-10y and 20+y Treasuries; TIPS; investment-grade credit; gold; broad commodities; REITs). Each is held at equal risk when its 12-month excess return is positive, otherwise T-bills, checked monthly.
- Q2 benchmark: the same sleeves at equal risk, always held.
- Mechanism: many weakly correlated trends.
- Expected: far below 0.76, because there is no short side, about 10 markets instead of 67, and the ETF era overlaps the weak 2010s. A plausible gain over the static mix is 0 to +0.2 (inference). comprehensive_v1's GTAA/DAA-style relatives showed +0.17 in-sample, with timing p 0.92.
- Data: run on a pre-ETF rebuild first (stocks 1926+, Treasuries rebuilt from yields 1962+, other sleeves where long series exist). Charge it as a replication of comprehensive_v1's tactical block.
- Kill: 0 or below on the rebuild, or negative in most decades.

3. VOLATILITY MANAGEMENT OF STOCKS, CREDIT AND EQUITY-HEAVY MIXES ONLY.
- Design: a 20-day EWMA anchor with 10/40/60/90-day neighbours and a 10% target, capped at 100% invested, so it can only reduce risk.
- Support: Harvey et al., including 60/40 going from 0.80 to 0.87-0.91; Cederburg et al.
- Expected: about 0 to +0.05 unlevered. Pre-declare drawdown and worst month as secondary outcomes.
- Pass: a gain above 0 across the whole plateau in both 1936-2004 and 2005-2022, or a pre-declared tail improvement at no Sharpe cost.
- Kill: 0 or below in 1936-2004; the verifier's rerun already suggests about +0.02. Then keep it only as a risk-control choice, never as an alpha claim.
- Never apply it to TLT/IEF/GLD/DBC.

4. TREASURY MATURITY BY CARRY.
- Design: hold whichever of SHY/IEF/TLT has the highest carry per unit of duration (yield plus roll-down minus the T-bill rate), checked monthly. Q2 compares it with the average-duration static mix.
- Support: Koijen et al., in-sample only. US Treasury carry scored 0.68 vs 0.57 passive, while slope carry lost to passive globally.
- Expected: +0.1 or less (inference).
- Data: returns rebuilt from FRED constant-maturity yields for 1962-2004, then the ETFs.
- Kill: 0 or below over 1962-2004.
- This is new to the lab, which has used the 10y-2y spread only to time stocks.

5. INTERNATIONAL GATE.
- Rerun anything that passes, unchanged, on the country ETFs (six from 1996, three from 2000) or on longer country indices.
- Require the same sign in a clear majority. Bongaerts et al.'s US vol-targeting gain shrank from +0.15 to +0.04 averaged across 10 markets (verifier note).

6. EVENT-STYLE CALENDAR EFFECTS (turn-of-month; pre-FOMC drift).
- No claim survived this review, and the verified calendar evidence is negative.
- But these events recur about 12 times a year, so one pre-registered test on 1926+ daily data has real power (inference).
- A turn-of-month test would also help read Thales' own start-day sweep, where only schedules starting near a month boundary were positive.
- Kill if the average return after publication is 0 or below net of 5 bp costs, which is what happened to the Monday effect.

Sources: https://www.aqr.com/Insights/Research/Journal-Article/A-Century-of-Evidence-on-Trend-Following-Investing, https://people.duke.edu/~charvey/Research/Published_Papers/P135_The_impact_of.pdf, https://www.lehigh.edu/~xuy219/research/COWY.pdf, https://www.nber.org/system/files/working_papers/w19325/w19325.pdf, https://escholarship.org/uc/item/2z02z6d9, /Users/dan/Documents/GitHub/thales/research/lab/README.md, /Users/dan/Documents/GitHub/thales/CLAUDE.md

</details>

## Caveats

COVERAGE: no claim survived verification on sector/factor rotation, pre-FOMC drift, turn-of-month or Halloween effects, McLean-Pontiff-style rates of post-publication decay, Harvey-Liu-Zhu or López de Prado / Bailey-Borwein-López de Prado-Zhu guidance, alpha-investing or alpha-spending, or reuse of a holdout. 'Time-series momentum: is it there?' (Huang et al.), Bongaerts et al., Kim-Tse-Wald and Hansen-Lunde-Nason appear only in verifier notes.

INFERENCE, NOT VERIFIED: three parts of this synthesis are its own reasoning from verified numbers plus the lab's code:
- the power calculations, which assume normal, serially independent differences and ρ ≈ 0.8;
- the analysis of the lab's error budget;
- the ranked run list and the expected gains for long-only versions.

INSTRUMENT MISMATCH: nearly all the positive evidence comes from long/short, futures-based, CRSP-index or levered portfolios. The lab is long-only, unlevered and ETF-based with 5 bp costs, so short legs and gains that depend on leverage do not transfer.

INTERESTED PARTIES: the trend, carry and CAPE papers come from AQR authors and the vol-targeting paper from Man AHL. The CAPE and vol-targeting papers do argue against popular uses of those ideas.

VERSIONS:
- Campbell-Thompson's simulations exist only in the 2005 NBER working paper.
- The Denk-Löffler figures are from the 2022 working paper; the published version splits at 1993/94 and its numbers differ slightly.
- The Sullivan et al. figures are from the 1998 working paper.
- The Koijen et al. figures are from the 2013 NBER version. Its headline moved between drafts, and the JFE 2018 version was not checked.

SPLIT VOTE: the weakening of trend following in 2010-16 passed only 2-1.

REFUTED AND NOT USED (7 claims). Refutation may mean wrong details rather than that the opposite is true.
- Campbell-Thompson restrictions make many predictors beat the historical mean (1-2).
- Long-only caps shrink the value of predictability (0-3).
- Sum-of-the-parts is the one approach that survives (1-2).
- 12-month momentum timing lifted Sharpe from 0.38 to 0.48 over 1900-2015 (0-3).
- Carry works as a single-asset timing signal (0-3).
- Specific Goyal-Welch-Zafirov statistics for the dividend-price ratio and the default and term spreads (1-2).
- Only 3 of 184 Goyal-Welch-Zafirov strategies beat all-equity (0-3).

DATA REUSE: pre-ETF history is new to the lab but not to the literature. The classic rules were partly designed on it, so long-history tests of those rules are biased upward. Only data a rule's authors never saw is clean.

TIME: every verified sample ends between 2011 and 2021. Nothing here covers 2022-2026, which is the lab's sealed holdout.

## Open questions

- How much of the long/short diversified trend and carry results survives in a long-only portfolio of about 10 ETFs, and what is the realistic timing gain over its static mix? No verified source tested a long-only version.
- How should the lab split its 5% error budget across a sequence of studies: a pre-declared confirmatory reserve, weighted Bonferroni, gatekeeping, or alpha-investing / online false-discovery-rate control? And how should old ledger rows be re-scored under the new scheme? The verified evidence did not cover this.
- Do effects that recur many times a year (turn-of-month, pre-FOMC drift) survive after publication and after 5 bp costs? And is the momentum sleeve's pattern of only near-month-boundary start days being positive a turn-of-month effect or schedule luck?
- Does the sum-of-the-parts equity forecast, the one exception in Dichtl et al.'s abstract and contested 1-2 in verification, hold up in 1994-2022 and outside the US, or did it fade like the other forecast combinations?

## Claims refuted in verification (not used)

- The classic valuation and yield-spread predictors keep failing. The paper runs every variable through one common monthly regression on the log equity premium, over the full available sample to 2021. In it, the dividend-price ratio has a Newey-West t of 0.71 and an out-of-sample R² of -0.06%, with forecasts restricted as Campbell-Thompson propose. The default-yield spread dfy (Baa minus Aaa yields, as defined in Goyal-Welch 2008) has t = 0.12 and an out-of-sample R² of -0.12%, and it underperformed all-equity in all four timing strategies. The term spread tms has t = 1.10 and an insignificant out-of-sample R² of +0.02%. Since the 2008 paper, out-of-sample performance of 13 of the original 17 predictors has deteriorated further, which the authors call worse than chance. The best pre-2008 survivors, the T-bill and long-term Treasury yields, owe most of their record to the 1974-75 oil-shock episode, and their timing strategies lost money in absolute terms.
- Statistical predictability almost never became timing value against buy-and-hold. Each of the 46 predictors was turned into 4 simple stock-versus-Treasury timing strategies (untilted or equity-tilted, with fixed or z-score-scaled bets), 184 strategies in all. Only 3 beat all-equity-all-the-time: accrul, gpce and gip, all annual predictors, all heavily equity-tilted, and each in only one strategy. Their t-statistics were weak, at 1.31 to 1.54. 84% of the strategies underperformed all-equity, and about one-third of those also lost money in absolute terms, doing worse than staying out of the market. The authors conclude that they would not tilt their own portfolios with much enthusiasm on these signals.
- Weak restrictions based on theory make many standard equity-premium predictors beat the historical-average forecast out of sample. The two restrictions are: the slope must have the sign theory expects (otherwise it is set to zero), and the forecast premium must be at least zero. Imposing steady-state valuation-model restrictions does even better, because it removes the need to estimate the mean from a short, noisy return sample. The gains are small but economically meaningful. In the working-paper tables (monthly S&P 500 total returns, forecasts evaluated from 1927 to 2003), the restrictions raised the number of predictors with a positive out-of-sample R-squared from 7 of 15 to 12 of 15, although two of those 12 are about zero. Provenance: the published abstract comes from the authors' Harvard DASH repository, because the OUP page shows no abstract. The table figures come from NBER Working Paper 11468 (June 2005, data through 2003).
- Under realistic long-only constraints, the portfolio value of return predictability shrinks sharply, because timing cannot pay when the cap on the equity weight binds. The test setup was: no shorting, equity weight between 0% and 150%, risk aversion of 3. With doubly restricted forecasts, single predictors added at most about 0.13% per month in utility, with the Lettau-Ludvigson consumption-wealth ratio (cay) the exception at 0.30-0.38% per month. 7 of 15 predictors (constant-variance estimate) and 4 of 15 (rolling 5-year variance estimate) produced utility losses. Transaction costs were excluded. The authors note that costs could offset modest timing gains, but also that about 10 bp per month (1.2% per year) would cover substantial costs. This applies directly to unlevered 0-100% ETF timing rules, where the upside from high forecasts is capped even more tightly.
- Only a few variants of Ferreira & Santa-Clara's (2011) sum-of-the-parts (SOP) approach survive. They deliver robust, statistically significant economic gains over the historical mean after the data-snooping correction and after transaction costs. SOP is an economically restricted method: it forecasts the return's components (dividend-price ratio, earnings growth, change in the price-earnings multiple) separately instead of fitting free predictive regressions. That description comes from FSC 2011, not from this abstract. The abstract credits 'only few' SOP-based strategies, not the SOP family as a whole, so SOP-style economically restricted forecasting is the one equity-premium timing approach here with support after correction.
- 12-month time-series momentum timing, built the same way (50–150% tilt), raised gross Sharpe only modestly. Over 1900–2015 it scored 0.48 against buy-and-hold's 0.38, and an equal-weight value+momentum blend scored 0.43. Over 1958–2015 the figures were 0.43 and 0.41 against 0.37. Even over 116 years these margins were statistically insignificant. The results are gross of costs, and momentum's higher turnover costs are not deducted. So a realistic Sharpe gain from single-market timing is about +0.05 to +0.10, which a century of data cannot confirm at conventional significance.
- Carry also works as a pure timing signal on one asset at a time. The rule goes long a security when its carry is positive and short when it is negative, equal-weighted within each asset class. It earned positive Sharpe ratios in every asset class (Table VII Panel A): equity-index futures 0.40, global 10Y bonds 0.65, 10Y-2Y slopes 0.72, US Treasuries 0.60, commodities 0.40, currencies 0.78, credit 0.64. Equities reach 0.72 when timed against a 5-year rolling mean of carry instead of zero. These are long/short futures portfolios, not long-only switches. US equity carry (expected dividend yield minus the short rate) averaged -1.4%/yr over 1988-2012 (Table I), so a long-only 'hold only when carry > 0' version would sit in cash whenever the short rate exceeds the expected dividend yield. The paper describes the Panel B benchmark two ways: the text says the 'sample mean', the table caption says a 5-year rolling mean. It is therefore unclear whether Panel B avoids using future information.

## Sources

- https://people.duke.edu/~charvey/Research/Published_Papers/P135_The_impact_of.pdf
- https://www.lehigh.edu/~xuy219/research/COWY.pdf
- https://www.aqr.com/Insights/Research/Journal-Article/A-Century-of-Evidence-on-Trend-Following-Investing
- https://academic.oup.com/rfs/article/37/11/3490/7749383
- https://academic.oup.com/rfs/article-abstract/21/4/1509/1567518
- https://academic.oup.com/raps/article-abstract/14/4/545/7643730
- https://www.sciencedirect.com/science/article/abs/pii/S0169207020300510
- https://www.aqr.com/library/journal-articles/market-timing-sin-a-little
- https://www.nber.org/system/files/working_papers/w19325/w19325.pdf
- https://escholarship.org/uc/item/2z02z6d9
- https://academic.oup.com/rfs/article/33/1/75/5494694
- https://pmc.ncbi.nlm.nih.gov/articles/PMC7525326/
- https://www.sciencedirect.com/science/article/pii/S1062940826000756
- https://www.tandfonline.com/doi/abs/10.2469/faj.v69.n4.4
- https://econjwatch.org/articles/revisiting-hypothesis-testing-with-the-sharpe-ratio
- https://www.ams.org/notices/201405/rnoti-p458.pdf
- https://people.duke.edu/~charvey/Research/Published_Papers/P143_False_and_missed.pdf
- https://rss.onlinelibrary.wiley.com/doi/abs/10.1111/j.1467-9868.2007.00643.x
- https://www.science.org/doi/10.1126/science.aaa9375
- https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3167017
- https://www.nber.org/papers/w21329
- https://people.duke.edu/~charvey/Research/Published_Papers/P138_A_backtesting_protocol.pdf
- https://blog.thinknewfound.com/2018/04/diversifying-the-what-how-and-when-of-trend-following/
- https://allocatortraining.com/wp-content/uploads/2023/06/A-Quantitative-Approach-to-Tactical-Asset-Allocation.pdf

## Proposal (pending owner decision — nothing below has been built or run)

The research's central point, checked independently: the lab cannot see
realistic timing effects on 2005-2022 ETF data. A +0.1 Sharpe timing gain
needs ~250 years of data for 80% power at a plain 5% test and ~920 years at
today's lab-wide bar (a new one-rule study must now reach p <= 3.7e-5, because
the lab-wide correction gives each of the 1,344 exploratory trials an equal
slice of 5%). So the next move is to BUY POWER, not to generate variants:

1. **Error budget, declared in advance.** Split the 5% family-wise budget into
   a confirmatory reserve (e.g. 2.5% across at most five pre-registered,
   literature-fixed rules, 0.5% each, one-sided z ~ 2.58) and an exploratory
   pool (2.5%, shared by all grids under the existing lab-wide scaling).
   Fixed and committed to the ledger before any confirmatory result exists.
   Past runs are all nulls under either scheme; nothing is rescued.
2. **Long-history data, owned like the ETF panel.** Ken French daily US
   market + T-bill from 1926 (the 3-factor daily file; the repo's 5-factor
   file starts 1963); Treasury returns rebuilt from FRED constant-maturity
   yields from 1962; optionally Ken French industry portfolios for sectors.
3. **Confirmatory tests, in order** (one literature-fixed anchor rule each,
   plus a few one-setting neighbours that only check for a plateau; pass =
   positive overall + same sign in >= 2 of 3 sub-periods + survives dropping
   its best episode + p under its pre-declared share; kill = the reverse,
   which closes the family on ETF data too):
   a. single-market trend (10-month MA; 12-month TSMOM; 1/3/12 blend) on US
      stocks 1926-2004 vs T-bills — settles "low power vs no effect";
   b. diversified long-only trend: 8-10 asset-class sleeves at equal risk,
      each held when its 12-month excess return > 0 else T-bills, monthly;
      Q2 = the same sleeves always held;
   c. volatility management of the equity/credit sleeve only (20-day EWMA,
      10% target, capped at 100%) — judged as risk control, never alpha;
   d. Treasury maturity by carry (SHY/IEF/TLT by yield + roll-down per unit
      of duration) from 1962;
   e. international replication gate for anything that passes;
   f. optional, owner's call: one pre-registered turn-of-month test (many
      events per year = real power; relates to the B5 pinned forward read,
      which it must not touch).
4. **Stop spending trials on:** macro/valuation timing of stocks (credit and
   term spreads, CAPE, yields, claims dashboards); broad calendar grids;
   volatility scaling of bonds/gold/commodities; further rotation, canary and
   trend variants scored only on 2005-2022.
