How High Should Sharpe Be After You Search Many Strategies?
Research-search assumptions
The exact calculation assumes independent strategy trials and independent Normal returns with zero true expected return. Real parameter variants are usually correlated, so the entered trial count should not be treated as an automatic estimate of effective search breadth.
Search-adjusted hurdle
The hurdle rises because the reported strategy is the winner
For one zero-edge strategy under this benchmark, annualized sample Sharpe is a scaled Student t statistic. Once a researcher tests N independent variants and reports the maximum, the relevant null distribution is no longer one strategy. It is the best-of-N distribution.
If F is the single-strategy Sharpe cumulative distribution, then P(max Sharpe ≤ x) = F(x)N. This calculator chooses the threshold x so that the probability of at least one null strategy exceeding it equals the entered family-wise error rate.
A one-year Sharpe of 2 can be weak evidence after a large search
With 252 daily observations and one pre-specified zero-edge strategy, the one-sided 5% Sharpe hurdle is about 1.65. With the same sample length but 100 independent trials, the 5% family-wise hurdle rises to about 3.32. The result is not a claim that every Sharpe below 3.32 is false. It is the exact cutoff for this deliberately simple null model and search design.
This complements the selection-bias benchmark, which shows the full best-of-N null distribution, and the multiple-testing calculator, which works directly in false-positive probabilities.
More observations lower the hurdle
Noise shrinks as the return sample grows. At a 5% family-wise rate and 100 independent trials, the annualized null hurdle is about 4.76 with 126 observations, 3.32 with 252, 1.90 with 756, and 1.47 with 1,260. The annualization convention is 252 periods per year.
Those values do not define universal minimum Sharpe ratios. Changing the sampling frequency, dependence structure, return distribution, search process, or alternative hypothesis changes the evidentiary problem.
Independence is the important boundary
Real strategy variants frequently share signals, parameters, markets, and observations. One hundred neighboring lookbacks rarely behave like one hundred independent hypotheses. Treating them as independent can overstate the effective search breadth, while counting only the few reported finalists can understate a much larger adaptive research process.
The benchmark therefore does not infer an effective trial count from correlated variants and does not replace Deflated Sharpe Ratio, Probability of Backtest Overfitting, walk-forward testing, or a documented research ledger. It isolates one exact reference case.
Reusable threshold grid
The published grid crosses 126, 252, 756, and 1,260 observations with 1, 5, 20, 100, and 500 independent strategy trials at a 5% family-wise error rate. Each row reports the equivalent Šidák per-trial tail probability and annualized Sharpe threshold.
Citation and reuse kit
This study is available for factual citation and reuse. Preserve the stated assumptions and limitation when they materially affect interpretation, and use the stable canonical URL rather than a temporary search or distribution link.
Preferred citation: Bailey, Lee. “How High Should Sharpe Be After You Search Many Strategies?” Grizzly Bulls, September 15, 2026. https://grizzlybulls.com/backtest-sharpe-significance-threshold
Research question: After testing many independent zero-edge strategies, how high must the winning annualized Sharpe be for the entire research search to cross a chosen family-wise false-positive threshold?
Key finding: At a 5% family-wise error rate, 252 daily observations and 100 independent zero-edge strategy trials require an annualized winning Sharpe of about 3.32, versus about 1.65 for one pre-specified strategy.
Core assumptions: Independent strategy trials; independent Normal returns; zero true expected return; one-sided family-wise tail probability; 252 annualization periods.
Interpretation boundary: Real strategy variants are usually correlated and adaptive research can make the effective search breadth difficult to count, so the benchmark does not infer an effective number of independent trials from a real parameter sweep.
Research context
Harvey, Liu, and Zhu argue that extensive factor search raises the significance hurdle for newly discovered return predictors rather than leaving the conventional threshold unchanged. Their empirical multiple-testing framework is richer than the independent-null benchmark here and explicitly addresses dependence among tests.
Bailey and López de Prado's Deflated Sharpe Ratio work likewise treats strategy selection and multiple testing as reasons not to evaluate a selected Sharpe as though it came from one pre-specified trial.
Harvey, Liu & Zhu: … and the Cross-Section of Expected Returns · Bailey & López de Prado: The Deflated Sharpe Ratio