Learn · Trading & Charts · Regimes, Systems and Statistics
Performance Statistics
Expectancy — win rate times average win, minus loss rate times average loss, in units of risk — is the number a method lives on, and a 35% win rate is profitable whenever winners average more than about 1.86 times the losers. The other statistics describe the ride (drawdown, Sharpe, Sortino, MAR), and the standard error says whether the record is long enough to tell an edge from luck.
Win rate and payoff: two halves of one number
The two statistics traders quote most — how often they win and how much they win when they do — are meaningless apart. Together they make **expectancy**: the average result per trade, in R. Expectancy = win rate × average win − loss rate × average loss. A method with a positive expectancy makes money over enough trades; one with a negative expectancy loses it, however often it wins. The same arithmetic gives the **break-even win rate** for any payoff ratio: 1 ÷ (1 + payoff), where payoff is the average win divided by the average loss. At a 1 : 1 payoff you need to win half your trades just to stand still; at 3 : 1 you need only 25%. That is why trend-followers live comfortably with win rates in the 30s and why some high-win-rate strategies quietly lose: a method that wins 80% of the time but loses five times as much when it is wrong has a break-even of 83%. The trade-off is structural, not a matter of skill. Tight profit targets raise the win rate and shrink the payoff; letting winners run does the opposite. A method is a choice of where on that curve to sit — and the only question is whether its expectancy, after costs, is positive. Win rate against payoff — Payoff 1 : 1: Break-even win rate 50% · Payoff 2 : 1: Break-even 33% · Payoff 2.6 : 1 (the prediction): Break-even 28% — at 35% the edge is +0.26R · Payoff 0.2 : 1 (win 80%, lose 5×): Break-even 83% — at 80% the edge is negative ←
The ride: drawdown, Sharpe, Sortino and MAR
Expectancy says where a method ends up on average. It says nothing about the path, and the path is what a trader has to survive. Four measures describe it. **Maximum drawdown** is the largest fall from a peak in the equity curve to a later trough, in R or in percent. It is the number that decides whether a trader abandons a method at the worst moment. **The Sharpe ratio** divides the average excess return by its standard deviation — return per unit of total volatility — and is usually annualised (multiplying a daily ratio by the square root of about 252 trading days). **The Sortino ratio** uses only downside deviation in the denominator, so a method is not penalised for volatile gains. **MAR** (also called the Calmar ratio over shorter windows) divides the compound annual return by the maximum drawdown: how much the method earns for each unit of its worst fall. Each measure answers a different question, and each can be gamed by choosing the window. A Sharpe ratio computed over a calm three years looks excellent until the first crisis; a maximum drawdown over a short record is a lower bound on what the method can do, not an estimate of it. • Max drawdown — the worst peak-to-trough fall; the number that makes people quit. • Sharpe — return per unit of total volatility; annualise a daily figure by about √252. • Sortino — return per unit of downside volatility only. • MAR / Calmar — compound return divided by the max drawdown. A short record’s maximum drawdown is the smallest the method has shown, not the largest it can produce. Plan for worse.
How many trades before you know anything
The statistic almost nobody computes is the one that decides whether the others mean anything: the **standard error** of the expectancy. Trade results are noisy — R-multiples routinely have a standard deviation of 1.5R or more — so the average of a small sample wanders far from the true edge. The standard error is the standard deviation divided by the square root of the number of trades, and a rough rule is that an edge is distinguishable from zero only when the average sits about two standard errors above it. The arithmetic is sobering. With a true edge of +0.2R and a standard deviation of 1.5R, 30 trades give a standard error of about 0.27R — larger than the edge itself. To put the average two standard errors clear of zero takes roughly (2 × 1.5 ÷ 0.2)² ≈ 225 trades. A year of a swing trader’s results can be pure noise in either direction, which is why a forward test is judged on a sample size fixed in advance (T22), and why a good month says almost nothing. The same logic runs the other way. A short losing streak is not evidence that an edge has disappeared; the Monte Carlo lab in T17 showed how long the streaks of a sound method run. The question is always whether the evidence has outgrown the noise. Edge +0.2R, spread 1.5R — 30 trades: Standard error ≈ 0.27R — larger than the edge · 100 trades: Standard error ≈ 0.15R — the edge is 1.3 standard errors · 225 trades: Standard error ≈ 0.10R — the edge is 2 standard errors ←
How uncertain a Sharpe ratio is
A Sharpe ratio looks precise and is not. Andrew Lo showed how to estimate its standard error, and the numbers are humbling. A strategy with a true annual Sharpe ratio of 1.0, measured on three years of monthly returns, has a standard error of about 0.59 — so the 95% range around the estimate runs from about −0.15 to about 2.15. Three years of a genuinely good strategy cannot be told apart from three years of a mediocre one. Ten years helps but does not settle it: the same strategy measured over a decade has a standard error of about 0.32 and a range of roughly 0.37 to 1.63. A rough rule of thumb follows from the arithmetic: the number of years needed for a Sharpe ratio to clear two standard errors is about (2 ÷ Sharpe)². A strategy with a Sharpe of 1.0 needs about four years; at 0.5 it needs about sixteen; at 0.3, the level of many real investment strategies, it needs more than forty. The consequence for reading anyone’s record — including your own — is that the statistics in this lesson need a sample before they mean anything, and the sample is usually longer than the record. Judge a short record by its process and its risk controls rather than its Sharpe ratio, compare it with the luck band, and be most suspicious of the results that look best. • Standard error of a Sharpe ratio (Lo): large for short records. • Sharpe 1.0 over 3 years of monthly data: 95% range ≈ −0.15 to 2.15. • Years to clear two standard errors ≈ (2 ÷ Sharpe)². • Judge short records by process and risk control, not by the ratio. How long a record has to be (to clear two standard errors) — Sharpe 1.0: About 4 years · Sharpe 0.5: About 16 years ← · Sharpe 0.3: More than 40 years
The statistics that decide, and the one that matters most
Four numbers describe a system well enough to make a decision. Expectancy in R, which is the average result per trade and is the only figure that answers “does this pay”. The break-even win rate, one over one plus the payoff ratio, which tells you how much room the observed win rate has — because a win rate without the payoff is uninterpretable. The maximum drawdown, which is what determines both the capital the system needs and whether it can be tolerated: a 20% drawdown on a strategy the trader abandons at 15% is a strategy with a 15% drawdown. And the risk-adjusted return, most often Sharpe, which divides the excess return by its volatility: (21.5 − 4.0) ÷ 14.8 = 1.18 on the worked example, with the MAR ratio — annual return over maximum drawdown, 21.5 ÷ 11.4 = 1.89 — as the version that speaks in units a trader can feel. Those four numbers are necessary and not sufficient, because every one of them was computed on the sample the rules were tuned against. The statistic that decides is the out-of-sample result, and the honest way to read it is as a *shrink* rather than a verdict. A system with +0.40R in sample and +0.07R out of sample has told you something precise: the marginal information in the rules is small relative to the noise, and the sample was too thin to distinguish an edge from a coincidence. The standard responses are to widen the sample in time rather than in trades-per-parameter, to cut the parameters to the ones a mechanism justifies, and to accept that a forward test of a hundred trades now takes months — which is the price of knowing. The last piece is the one most often skipped: the conditions inside the test window. A trend-following system tested over a decade that contained one long bull market and one short crash has been tested on a handful of independent episodes rather than thousands of independent trades, because trades within a trend are not independent observations. That is why the write-up should include what regimes the sample contained, how the system behaved in the worst of them, and what it would have required in position size to ruin the account during the worst losing run. The complete statement is short and unglamorous: here are the rules, here is the sample and its conditions, here is the expectancy and its break-even, here is the worst drawdown, and here is the out-of-sample result with its error bar. • Expectancy in R, with the break-even win rate beside it — a win rate alone is uninterpretable. • Maximum drawdown as a capital requirement and a tolerance test, because the tradable drawdown is the one the trader can hold. • Sharpe for risk-adjusted return and MAR for a version measureable in the account currency. • The out-of-sample result as the deciding number, read as a shrink rather than a verdict. • The regimes inside the sample, because trades inside one trend are not independent observations. The most tempting error is to re-optimise after an out-of-sample failure, which converts the out-of-sample period into part of the search and leaves the system with no untouched data at all. If the rules must change, the honest sequence is to freeze them, note why they changed against a mechanism rather than a result, and hold back a new period — and to record the number of attempts, because a tenth variation that works out of sample is the tenth attempt rather than the first.
What you'll practise
A method wins 40% of the time; winners average +2R and losers −1R. What is its expectancy?
40 XP in the app · multi select
Sources
- Expectancy and R-multiplesVan Tharp, “Trade Your Way to Financial Freedom”
- The Sharpe ratioSharpe, “The Sharpe Ratio” (Journal of Portfolio Management, 1994)
- Statistical significance of trading resultsAronson, “Evidence-Based Technical Analysis” (2007)
- Sortino and downside deviationSortino & Price, “Performance Measurement in a Downside Risk Framework” (Journal of Investing, 1994)
- Overfitting, data snooping and the deflated Sharpe ratioBailey & López de Prado, “The Deflated Sharpe Ratio” (Journal of Portfolio Management)
- Look-ahead bias, survivorship and the cost of a backtest that liesChan, “Algorithmic Trading” and “Quantitative Trading”
- Expectancy, R multiples and the risk of ruinVan Tharp, “Trade Your Way to Financial Freedom”
Learn content is for education only — not individualized financial advice, a recommendation, or a solicitation to buy or sell any security. Options involve substantial risk. Examples are simplified and historical patterns never guarantee future results.