Backtest Win Rate Confidence Calculator

A 60% win rate from 100 trades and a 60% win rate from 1,000 trades are not equally informative. This tool measures that sampling uncertainty instead of treating the headline percentage as exact.
60 wins / 100 trades95% Wilson interval: 50.20% to 69.06%
600 wins / 1,000 trades95% Wilson interval: 56.93% to 62.99%

Backtest observations

The payoff ratio is used only for the simple pre-cost break-even comparison. A value of 1 means the average winner and average loser have equal absolute size.

What the sample supports

60%Observed win rate
50.2%–69.06%95% Wilson interval
18.86 ppTotal interval width
50%Pre-cost break-even win rate at the entered payoff ratio
10 ppObserved margin above break-even
0.2 ppLower-bound margin above break-even

Why the headline win rate is not enough

A backtest win rate is a sample proportion. With only a finite number of trades, the observed percentage moves around even if the underlying process is unchanged. The Wilson interval expresses that sampling uncertainty without pretending the observed percentage is known exactly.

The sample-size effect is large. Sixty wins in 100 trades produces the same 60% headline rate as 600 wins in 1,000 trades, but the second interval is much tighter because it contains ten times as many observations.

Break-even depends on payoff size too

Win rate alone does not determine profitability. If the average winner is R times the average loser, the simple pre-cost break-even win rate is 1 ÷ (1 + R). At a 1:1 payoff ratio, break-even is 50%. At 2:1, it is about 33.33%.

The lower-bound comparison is intentionally conservative, but it is not a profitability guarantee. Trading costs, changing payoff distributions, tail losses, and dependence between trades can all invalidate the simple comparison.

What the confidence interval does not mean

A frequentist 95% confidence interval is not a statement that there is a 95% probability the true win rate lies inside this one realized interval. It describes the long-run coverage of the interval procedure under its assumptions.

The calculation also treats the trade outcomes as Bernoulli observations from a stable process. Regime changes, overlapping signals, correlated positions, walk-forward selection, parameter search, and repeated strategy mining can make the effective evidence much weaker than the raw trade count suggests.

Use it with the rest of the backtest evidence

This tool isolates sampling uncertainty in win rate. The losing-streak calculator addresses sequence risk, while the transaction-cost tool stress-tests execution friction. The algorithmic trading guide covers the broader research, validation, execution, and monitoring process.

Download the 95% sensitivity grid as CSV · JSON

Statistical method

The interval uses the Wilson score method for a binomial proportion. NIST describes the Wilson interval as the algebraic counterpart to inverting the large-sample proportion test and documents its bounds within the 0-to-1 range.

NIST: Wilson confidence limits for a binomial proportion