Deflated Sharpe Ratio Calculator

A winning backtest should not be judged against a zero hurdle if it was selected from a broad research search. Deflated Sharpe Ratio raises the benchmark to reflect the search itself, then asks how strongly the chosen Sharpe clears it.
Observed annualized Sharpe 2.0252 daily observations
100 trials with 0.50 Sharpe dispersionExpected maximum search benchmark: about 1.27

Selected backtest and research search

Trial mean and dispersion must use the same annualized Sharpe convention as the selected result. Pearson kurtosis is 3 for a Normal distribution.

Search-deflated evidence

76.741%Deflated Sharpe probability
1.265Expected maximum Sharpe benchmark from the entered search
0.735Selected Sharpe minus expected search maximum
0.73Search-deflated z-score

Why Deflated Sharpe Ratio is different from ordinary PSR

The Probabilistic Sharpe Ratio asks whether one observed Sharpe exceeds a chosen benchmark after accounting for finite sample length, skewness, and kurtosis. That is appropriate when the strategy was genuinely pre-specified.

Deflated Sharpe Ratio adds the research search. Instead of comparing the selected winner with an arbitrary zero benchmark, it estimates the Sharpe level that the strongest candidate could reach simply because many trials were considered. The selected result is then evaluated against that search-aware benchmark using the same non-Normal Sharpe uncertainty structure.

How the search benchmark is constructed

The Bailey and López de Prado formulation approximates the expected maximum Sharpe across the entered number of trials using the mean and standard deviation of Sharpe ratios across the research search. The approximation uses two upper Normal quantiles combined with the Euler-Mascheroni constant.

That benchmark rises when more trials are considered or when Sharpe outcomes vary more widely across the search. A selected Sharpe that looks impressive against zero can therefore carry much less evidence once the larger family of attempted results is acknowledged.

Trial count is not automatically effective search breadth

The raw number of tested variants can be misleading when strategies are strongly correlated, nested, or generated adaptively. Entering 500 nearly identical parameter settings as though they were 500 independent discoveries can overstate the deflation hurdle.

Our correlated strategy-search benchmark shows why dependence among tested variants matters. This calculator does not infer an effective number of independent trials from a correlation matrix or research log. That remains a separate modeling problem.

What DSR does not repair

Deflating Sharpe does not repair look-ahead bias, survivorship bias, data leakage, unrealistic fills, hidden costs, serial dependence, regime instability, or a contaminated holdout. The autocorrelation calculator covers one important serial-dependence sensitivity, while transaction-cost analysis addresses execution friction.

DSR is strongest when the researcher can honestly describe the family of trials that produced the winner. Undisclosed discarded tests make the entered search look narrower than it really was.

Reusable sensitivity data

The downloadable grid holds 252 daily observations, zero skew, Pearson kurtosis 3, trial mean Sharpe 0, and trial Sharpe standard deviation 0.50. It varies the selected annualized Sharpe from 1.0 to 2.5 and research breadth from 1 to 500 trials.

Download CSV · Download JSON

Method sources