Backtest Win Rate Confidence Calculator
Backtest observations
The payoff ratio is used only for the simple pre-cost break-even comparison. A value of 1 means the average winner and average loser have equal absolute size.
What the sample supports
Why the headline win rate is not enough
A backtest win rate is a sample proportion. With only a finite number of trades, the observed percentage moves around even if the underlying process is unchanged. The Wilson interval expresses that sampling uncertainty without pretending the observed percentage is known exactly.
The sample-size effect is large. Sixty wins in 100 trades produces the same 60% headline rate as 600 wins in 1,000 trades, but the second interval is much tighter because it contains ten times as many observations.
Break-even depends on payoff size too
Win rate alone does not determine profitability. If the average winner is R times the average loser, the simple pre-cost break-even win rate is 1 ÷ (1 + R). At a 1:1 payoff ratio, break-even is 50%. At 2:1, it is about 33.33%.
The lower-bound comparison is intentionally conservative, but it is not a profitability guarantee. Trading costs, changing payoff distributions, tail losses, and dependence between trades can all invalidate the simple comparison.
What the confidence interval does not mean
A frequentist 95% confidence interval is not a statement that there is a 95% probability the true win rate lies inside this one realized interval. It describes the long-run coverage of the interval procedure under its assumptions.
The calculation also treats the trade outcomes as Bernoulli observations from a stable process. Regime changes, overlapping signals, correlated positions, walk-forward selection, parameter search, and repeated strategy mining can make the effective evidence much weaker than the raw trade count suggests.
Use it with the rest of the backtest evidence
This tool isolates sampling uncertainty in win rate. The losing-streak calculator addresses sequence risk, while the transaction-cost tool stress-tests execution friction. The algorithmic trading guide covers the broader research, validation, execution, and monitoring process.
Statistical method
The interval uses the Wilson score method for a binomial proportion. NIST describes the Wilson interval as the algebraic counterpart to inverting the large-sample proportion test and documents its bounds within the 0-to-1 range.