Backtest Multiple Testing Calculator

Testing more strategy variants creates more chances to find a result that looks significant by luck. This tool makes that search-cost arithmetic visible instead of treating every reported p-value as if it came from one pre-specified test.
20 tests at 5% each64.15% chance of at least one false positive if the tests are independent true nulls
100 tests at 5% each99.41% chance of at least one false positive under the same independence assumption

Search assumptions

A “test” can be a strategy, parameter set, filter, market, lookback, threshold, or other hypothesis examined before the winner is selected. The true search count can be much larger than the number eventually reported.

Search-cost result

64.151%Chance of at least one false positive if all tests are independent true nulls
35.849%Chance none of those independent null tests crosses the entered threshold
1Expected false positives across the entered tests
0.25%Bonferroni per-test threshold for the target family-wise rate
0.256%Šidák per-test threshold under independence

Why strategy mining changes the evidence

If one valid null test has a 5% chance of producing a false positive, repeating that test logic across many independent no-edge variants creates many opportunities for luck to win. Under independence, the probability of at least one false positive is 1 − (1 − α)m, where α is the per-test false-positive rate and m is the number of tests.

That is why 20 independent 5% tests produce a 64.15% family-wise chance of at least one false positive, even though each individual threshold still says 5%.

Bonferroni and Šidák answer different assumptions

The Bonferroni threshold divides the target family-wise error rate by the number of comparisons. It follows from the Bonferroni inequality and does not require the tests to be independent, although it can be conservative.

The Šidák threshold is slightly less restrictive because it solves the independent-test equation directly. It should not be presented as exact when strategy variants are dependent, which is common when they share data, signals, parameters, or positions.

A correction does not repair a mined backtest

Multiple-testing arithmetic is only one layer of backtest validation. It does not repair look-ahead bias, survivorship bias, repeated dataset reuse, undocumented discarded variants, changing market regimes, poor execution assumptions, or a research process that keeps searching until something works.

The effective hypothesis family can also be difficult to count. If a researcher tried hundreds of thresholds, indicators, universes, and sample windows but reports only five finalists, using five as the test count understates the search process.

Use it with the rest of the backtest evidence

The win-rate confidence tool measures sampling uncertainty in one reported win rate. The transaction-cost tool measures implementation drag, while the losing-streak tool isolates sequence risk. The algorithmic trading guide covers the broader validation process.

Download the sensitivity grid as CSV · JSON

Statistical method

NIST documents the Bonferroni general inequality and its use for preserving an overall confidence or error budget across a finite set of comparisons. The independent-test family-wise and Šidák calculations shown here are algebraic consequences of multiplying independent non-rejection probabilities.

NIST: Bonferroni's method