Backtest Multiple Testing Calculator
Search assumptions
A “test” can be a strategy, parameter set, filter, market, lookback, threshold, or other hypothesis examined before the winner is selected. The true search count can be much larger than the number eventually reported.
Search-cost result
Why strategy mining changes the evidence
If one valid null test has a 5% chance of producing a false positive, repeating that test logic across many independent no-edge variants creates many opportunities for luck to win. Under independence, the probability of at least one false positive is 1 − (1 − α)m, where α is the per-test false-positive rate and m is the number of tests.
That is why 20 independent 5% tests produce a 64.15% family-wise chance of at least one false positive, even though each individual threshold still says 5%.
Bonferroni and Šidák answer different assumptions
The Bonferroni threshold divides the target family-wise error rate by the number of comparisons. It follows from the Bonferroni inequality and does not require the tests to be independent, although it can be conservative.
The Šidák threshold is slightly less restrictive because it solves the independent-test equation directly. It should not be presented as exact when strategy variants are dependent, which is common when they share data, signals, parameters, or positions.
A correction does not repair a mined backtest
Multiple-testing arithmetic is only one layer of backtest validation. It does not repair look-ahead bias, survivorship bias, repeated dataset reuse, undocumented discarded variants, changing market regimes, poor execution assumptions, or a research process that keeps searching until something works.
The effective hypothesis family can also be difficult to count. If a researcher tried hundreds of thresholds, indicators, universes, and sample windows but reports only five finalists, using five as the test count understates the search process.
Use it with the rest of the backtest evidence
The win-rate confidence tool measures sampling uncertainty in one reported win rate. The transaction-cost tool measures implementation drag, while the losing-streak tool isolates sequence risk. The algorithmic trading guide covers the broader validation process.
Statistical method
NIST documents the Bonferroni general inequality and its use for preserving an overall confidence or error budget across a finite set of comparisons. The independent-test family-wise and Šidák calculations shown here are algebraic consequences of multiplying independent non-rejection probabilities.