ClearViewLesson libraryWhat's new

Learn · Trading & Charts · The Plan and the Review

Systems, Backtests and the Statistics That Judge Them

40 min read

A backtest measures the search as much as the market: every parameter tuned on the same sample is a degree of freedom spent, so the only statistics that count are the ones computed on data the rules never saw — and the arithmetic of expectancy, drawdown and the out-of-sample collapse is what turns a promising curve into a decision.

Specify first, then look: the five ways a backtest lies

A backtest is an experiment, and like any experiment it is only as good as the constraints placed on it before the data is seen. The strongest constraint is a written specification: the entry trigger, the exit and stop rule, the filters, the position sizing rule and the instruments — all of them fixed in text before the first test is run. Everything that follows is measured against that document. Without it, there is no way to tell afterwards whether the rules came from a hypothesis or from the search, and that distinction is the entire difference between a system and a curve. Five biases account for most of the lies. *Overfitting* is the first: every parameter tuned against the same sample is a degree of freedom spent, and with enough of them a rule set can be fitted to noise — which is why the count of trades per optimised parameter is a diagnostic, and why ten is not enough. *Look-ahead bias* is the second: using information that was not available at the time, such as a close that had not yet printed, a restated figure, or an index membership decided later. Its signature is an implausibly smooth equity curve, and its cause is usually a data table that was as-of-today rather than as-of-then. *Survivorship bias* is the third: testing on today’s index members means every company that went bankrupt has been quietly removed, which flatters every strategy that buys weakness. *Costs* are the fourth, and they are the most common reason a beautiful curve is untradeable: spread, slippage and commission have to be modelled at the size the strategy will trade, because a strategy that scalps a fraction of a percent is a levered bet on execution quality. The fifth is *data snooping*, which is subtler than the others because it leaves no trace in the code. Testing twenty variations and reporting the best one, or trying a strategy on six instruments and trading the two that worked, converts a search into a result. The formal fix is uncomfortable: the reported significance has to be deflated for the number of attempts, and the practical fix is to pre-register the rules and the test, then count the attempts honestly. A related version applies to the sample itself — a system that only ever saw a rising market has no information about its behaviour in the other regime, which is why the market conditions inside the test window matter as much as its length. Five ways the curve lies — Overfitting: Six parameters on sixty trades — ten observations per degree of freedom ← · Look-ahead: Data assembled as-of-today, so the test sees information that did not exist · Survivorship: Today’s index membership, so the failures have been deleted · Costs: A scalping edge that vanishes once the spread is modelled at size ← · Data snooping: Twenty variations tried, the best one reported — significance deflated accordingly The order matters: a written specification comes first because none of these biases can be diagnosed afterwards from the results alone. A smooth curve is consistent with both a good system and a leaky test.

The statistics that decide, and the one that matters most

Four numbers describe a system well enough to make a decision. Expectancy in R, which is the average result per trade and is the only figure that answers “does this pay”. The break-even win rate, one over one plus the payoff ratio, which tells you how much room the observed win rate has — because a win rate without the payoff is uninterpretable. The maximum drawdown, which is what determines both the capital the system needs and whether it can be tolerated: a 20% drawdown on a strategy the trader abandons at 15% is a strategy with a 15% drawdown. And the risk-adjusted return, most often Sharpe, which divides the excess return by its volatility: (21.5 − 4.0) ÷ 14.8 = 1.18 on the worked example, with the MAR ratio — annual return over maximum drawdown, 21.5 ÷ 11.4 = 1.89 — as the version that speaks in units a trader can feel. Those four numbers are necessary and not sufficient, because every one of them was computed on the sample the rules were tuned against. The statistic that decides is the out-of-sample result, and the honest way to read it is as a *shrink* rather than a verdict. A system with +0.40R in sample and +0.07R out of sample has told you something precise: the marginal information in the rules is small relative to the noise, and the sample was too thin to distinguish an edge from a coincidence. The standard responses are to widen the sample in time rather than in trades-per-parameter, to cut the parameters to the ones a mechanism justifies, and to accept that a forward test of a hundred trades now takes months — which is the price of knowing. The last piece is the one most often skipped: the conditions inside the test window. A trend-following system tested over a decade that contained one long bull market and one short crash has been tested on a handful of independent episodes rather than thousands of independent trades, because trades within a trend are not independent observations. That is why the write-up should include what regimes the sample contained, how the system behaved in the worst of them, and what it would have required in position size to ruin the account during the worst losing run. The complete statement is short and unglamorous: here are the rules, here is the sample and its conditions, here is the expectancy and its break-even, here is the worst drawdown, and here is the out-of-sample result with its error bar. • Expectancy in R, with the break-even win rate beside it — a win rate alone is uninterpretable. • Maximum drawdown as a capital requirement and a tolerance test, because the tradable drawdown is the one the trader can hold. • Sharpe for risk-adjusted return and MAR for a version measureable in the account currency. • The out-of-sample result as the deciding number, read as a shrink rather than a verdict. • The regimes inside the sample, because trades inside one trend are not independent observations. The most tempting error is to re-optimise after an out-of-sample failure, which converts the out-of-sample period into part of the search and leaves the system with no untouched data at all. If the rules must change, the honest sequence is to freeze them, note why they changed against a mechanism rather than a result, and hold back a new period — and to record the number of attempts, because a tenth variation that works out of sample is the tenth attempt rather than the first.

How to split the data honestly

Every backtest has to decide what counts as the data it learned from and what counts as the data it is judged on, and the choice that decision exposes is the one the five ways a curve lies does not cover: the *number of times* the data was split. A single holdout is a test. Ten holdouts, of which one is reported, is a search with a witness. The simplest split is a single cutoff — build on everything before a date, test on everything after. The proportion is a trade-off rather than a rule: a longer training window gives more confidence in the parameters and leaves less evidence to test them; a shorter one does the reverse. A reasonable default is to hold out something in the region of a fifth to a third of the sample, chosen by date rather than at random, because the point of a holdout in trading is to test whether the method survives a period it did not see. The stronger version is **walk-forward** testing. Instead of one split, the window rolls: fit on the first two years, trade the next six months, then advance the whole window and repeat. What comes out is a sequence of out-of-sample results, which is far more informative than one number because it shows whether the method works everywhere or only in the single period that happened to be held back. It also exposes the stability of the parameters, since a method whose optimum drifts wildly from window to window is telling you it is fitting noise. Two cautions follow from taking this seriously. The first is that any decision you make *after* seeing the out-of-sample result — a small tweak, a filter, a change of universe — consumes that holdout permanently. It is no longer out of sample; it is now part of the training data, and the next honest test needs new data or a new period. This is why serious work keeps a final holdout untouched and resists the temptation to look at it until the specification is frozen. The second is that overlapping positions make observations less independent than they appear: a method holding ten correlated positions for a month does not have ten independent outcomes, and a walk-forward test with concurrency needs the folds purged of overlap, or the same market move gets counted as several successful tests. The practical form is short enough to write down: one clean holdout set aside and not looked at, walk-forward windows for everything else, and a rule that any change after testing consumes the test and requires a new one. Methods that fail that regime are not necessarily bad; they are unproven, which is a different claim. • One clean holdout, split by date, is the test; everything else is training. • Walk-forward windows show whether a method works everywhere or once. • Every decision made after seeing the holdout consumes it permanently. • Purge overlapping positions, or one market move will be counted as several tests. The tell that a system was searched rather than tested is the description of the process: a single split, a small tweak, a re-test, a filter added, and a result reported from the last attempt. The number of attempts is not on the card, and it is the number that decides how much the result is worth.

What you'll practise

A system wins 24 of 60 trades at +2.5R and loses the other 36 at −1R. What is its expectancy?

50 XP in the app · multi select

Sources

Practise this in the app →

Learn content is for education only — not individualized financial advice, a recommendation, or a solicitation to buy or sell any security. Options involve substantial risk. Examples are simplified and historical patterns never guarantee future results.