Probability of Backtest Overfitting Calculator

PBO asks a different question from a Sharpe significance test: when the research process selects the best strategy in-sample, how often does that same winner fall into the worse half of the strategy set out of sample?
One winning Sharpe is not enoughPBO needs the full aligned return matrix for all strategy configurations considered, including the losers you would normally discard.
96 × 8 zero-edge exampleWith 8 CSCV blocks, the included deterministic example produces 70 balanced splits and PBO of about 45.7%.

CSCV inputs

Rows are aligned return observations and columns are strategy configurations. An optional header row is allowed. The observation count must divide evenly into the chosen even number of partitions.

Selection-process robustness

45.71%Probability of Backtest Overfitting
70Balanced CSCV splits
55.56%Median OOS percentile rank of the IS winner
-0Mean OOS Sharpe of the IS-selected winner
42.86%Splits where the IS winner has negative OOS Sharpe
0Splits with an exact in-sample winner tie

PBO is the share of symmetric splits where the in-sample winner's OOS rank logit is at or below zero, meaning it lands at or below the OOS median.

How CSCV estimates backtest overfitting

Bailey, Borwein, López de Prado, and Zhu proposed Combinatorially Symmetric Cross-Validation specifically for investment backtests. The method divides the full history into an even number of contiguous blocks, then enumerates every way to assign half the blocks to an in-sample set and the complementary half to an out-of-sample set.

For every balanced split, this implementation computes annualized Sharpe for every strategy, selects the strongest in-sample configuration, and then asks where that exact configuration ranks among all strategies out of sample. With an ascending OOS rank scaled by N + 1, larger relative rank means better OOS performance. The logit is ln(omega / (1 - omega)). A logit at or below zero means the in-sample winner finished at or below the OOS median.

The final PBO is simply the fraction of those symmetric splits that land on that overfit side of the rank distribution. PBO is therefore a property of the selection procedure and strategy library, not a probability that one chosen strategy will lose money.

Why PBO needs the whole strategy matrix

A single reported Sharpe cannot tell you whether the research process repeatedly selected winners that failed to generalize. PBO needs aligned observations for all configurations that participated in the search. Leaving out discarded variants changes the comparison set and can make the research process look more stable than it was.

This is also why PBO complements rather than replaces the Deflated Sharpe Ratio. DSR asks how strongly one selected Sharpe clears a search-aware benchmark. PBO asks whether the selection rule itself repeatedly chooses configurations that rank poorly on complementary data.

Important boundaries

  • CSCV reuses the same historical sample in many symmetric combinations. It is a backtest-overfitting diagnostic, not a substitute for genuinely new live data.
  • The return matrix should include the real configuration family explored by the researcher. A cherry-picked subset weakens the interpretation.
  • This implementation uses Sharpe ratio as the per-split selection statistic. Different objective functions define a different selection process.
  • Exact in-sample winner ties are broken deterministically by column order and reported separately so ties cannot remain invisible.
  • PBO does not repair look-ahead bias, survivorship bias, bad point-in-time data, unrealistic costs, or serial dependence inside each strategy's returns.

For temporal dependence inside one strategy, use the autocorrelation check. For dependence across strategy variants, see the correlated-search benchmark.

Reusable example data

The public example is a deterministic 96-observation × 8-strategy zero-edge Normal matrix generated with seed 35,035. It is included to make the calculator and implementation independently reproducible, not as a claim that 45.7% is a universal null PBO.

Download example matrix CSV · Download example result JSON

Method sources