The selected Sharpe can collapse without anything changing in the market
Start with 100 independent strategies. Give each strategy 252 independent Normal daily returns with a true mean of zero, calculate annualized sample Sharpe, and select the best in-sample result. Then evaluate that selected strategy on a new independent 252-day holdout generated from the same unchanged zero-edge process.
The median selected in-sample Sharpe is 2.48. The median out-of-sample Sharpe is still 0.00. Only 0.69% of independent one-year holdouts match or exceed the median selected in-sample Sharpe.
20 strategy trials
Median selected in-sample Sharpe: 1.83
One-year holdout matches it: 3.41%
100 strategy trials
Median selected in-sample Sharpe: 2.48
One-year holdout matches it: 0.69%
500 strategy trials
Median selected in-sample Sharpe: 3.02
One-year holdout matches it: 0.139%
Search does not create future edge
Under this zero-edge null, the selected strategy has a 50% chance of a positive holdout Sharpe no matter how many in-sample variants were searched.
Why selection inflates the past but not the independent holdout
In sample, the researcher is not looking at one pre-specified strategy. The researcher is taking the maximum across many noisy estimates. That selection step moves the winner distribution upward even though every candidate has zero true expected return. The selection-bias benchmark calculates that winner distribution directly.
The holdout is different. In this benchmark, the new returns are independent of the data used to choose the winner. Conditioning on a strategy having won the in-sample search therefore does not change its zero-edge holdout distribution. Its median annualized Sharpe returns to zero.
A longer holdout makes large null Sharpes less likely
Search size does not affect the selected strategy's holdout distribution in this null benchmark, but holdout length does. With 126 out-of-sample daily observations, a zero-edge strategy has a 7.99% chance of posting an annualized Sharpe of at least 2 by sampling noise. With 252 observations that falls to 2.33%. With 756 observations it falls to 0.028%.
That is not a rule that every strategy needs three years of holdout data. It is a transparent illustration of the precision tradeoff: a longer independent sample narrows the range of large Sharpe ratios that pure noise can produce.
This is not proof that a simple train/test split solves backtest overfitting
The benchmark deliberately gives the researcher a pristine holdout that is never used during strategy selection. Real research rarely stays that clean. Researchers can inspect a holdout, revise the model, try another specification, reuse the same validation period, change the universe, or keep iterating until the apparent out-of-sample evidence also becomes part of the search.
Bailey, Borwein, López de Prado, and Zhu developed the Probability of Backtest Overfitting framework precisely because ordinary holdout logic can be unreliable when the research process repeatedly optimizes investment backtests. Their work uses combinatorially symmetric cross-validation to study overfitting across strategy configurations rather than assuming one untouched split resolves the problem.
What the null benchmark does not model
The calculations assume independent strategy trials in sample, independent Normal returns, zero true expected return, and a genuinely untouched independent holdout. Real strategy variants are usually correlated. Real returns can be serially dependent, skewed, fat-tailed, heteroskedastic, capacity constrained, and regime dependent. Real research can also leak information across train and validation periods.
The study therefore does not estimate the expected live decay of a real strategy and does not say that any particular observed Sharpe is false. It isolates one mechanism: selection can inflate the reported past even when the future data-generating process has not deteriorated at all.
Use the multiple-testing tool for family-wise false-positive arithmetic, the Probabilistic Sharpe Ratio tool for finite-sample Sharpe uncertainty, and the backtest validation stack for the wider research process.
Reusable benchmark data
The published grid fixes the in-sample search at 252 daily observations and crosses 1, 5, 20, 100, and 500 independent trials with 126-, 252-, and 756-observation holdouts. Each row reports the median selected in-sample Sharpe, the holdout median, the chance of positive holdout Sharpe, the chance of holdout Sharpe at least 1 or 2, and the chance the holdout matches or exceeds the selected in-sample median.
Citation and reuse kit
This study is available for factual citation and reuse. Preserve the stated assumptions and limitation when they materially affect interpretation, and use the stable canonical URL rather than a temporary search or distribution link.
Preferred citation: Bailey, Lee. “What Happens Out of Sample After You Select the Best Backtest?” Grizzly Bulls, September 15, 2026. https://grizzlybulls.com/backtest-out-of-sample-decay
Research question: After selecting the strongest in-sample backtest from many zero-edge candidates, what does that selected strategy look like on a genuinely independent holdout when the underlying process has not changed?
Key finding: After selecting the best of 100 one-year zero-edge backtests, the median selected in-sample Sharpe is about 2.48 while the independent one-year holdout median is 0; only about 0.69% of holdouts match or exceed that selected in-sample median.
Core assumptions: Independent strategy trials in sample; independent Normal returns; zero true expected return; genuinely untouched independent holdout.
Interpretation boundary: Reusing the holdout for further model selection makes it part of the research search, so this benchmark does not claim one train/test split solves backtest overfitting.
Method sources and related research
- Bailey, Borwein, López de Prado & Zhu, The Probability of Backtest Overfitting, on strategy selection, performance degradation, and combinatorially symmetric cross-validation.
- Bailey, Borwein, López de Prado & Zhu, Pseudo-Mathematics and Financial Charlatanism, on how backtest optimization can manufacture strong in-sample performance and poor out-of-sample behavior.