ClearViewLesson libraryWhat's new

Learn · Risk & Sizing · The Tail and the Cost

Return per Unit of Risk

30 min read

A return is only half a result, and the other half is what it cost to get: Sharpe divides excess return by volatility (0.51 against 2.27 for the two strategies here), Calmar divides compound return by maximum drawdown (0.62 against 1.43), and Sortino counts only the downside. None is a perfect measure — each hides something — but comparing bare returns across strategies with different drawdowns is the mistake they exist to prevent, and the choice between the two depends on whether the household can hold the deeper path.

Three measures, three questions

**Sharpe** divides excess return by the standard deviation of returns: for A, (12 − 4) ÷ 15.62 = 0.51, and for B, (10 − 4) ÷ 2.65 = 2.27. It answers "how much return per unit of typical variation?", it is the standard comparison, and it has a specific blind spot — it treats upside and downside variation identically, so a strategy with large positive outliers is penalised for them. **Sortino** divides excess return by downside deviation only, which answers the question a household is actually asking: how much return per unit of bad months? It needs a target return, usually zero or the risk-free rate, and it is unhelpful for a strategy with no losing periods — the denominator is zero and the ratio is infinite, which is a reminder that these are ratios rather than truths. **Calmar** divides the compound return by the maximum drawdown: 11.23 ÷ 18 = 0.62 for A and 9.98 ÷ 7 = 1.43 for B. It answers "how much annual return per point of worst fall?", which is the measure that maps most directly onto the constraint from R2 — the depth of the hole a household has to sit through. It is also the most unstable of the three, because it depends on a single observation in the record, the worst one. Used together they triangulate, and the disagreements are informative. A strategy can have a high Sharpe and a low Calmar (a smooth series with one terrible quarter), or a high Calmar and a low Sharpe (a violent but shallow-riding series). The mistake to stop making is the simplest one: comparing two strategies on average return alone, when one of them costs three times the drawdown to deliver a point more. The two strategies, four ways — A: mean 12%, σ 15.62%, max drawdown 18%: Sharpe 0.51 · Calmar 0.62 · B: mean 10%, σ 2.65%, max drawdown 7%: Sharpe 2.27 · Calmar 1.43 ← · What A has more of: both return and volatility · What the comparison needs: the return and the path Sharpe depends on the risk-free rate, so it is a comparison rather than an absolute property: raise the risk-free rate to 8% and A falls to 0.26 while B falls to 0.76, and the ranking stands.

The measure a household needs, and its limits

The household constraint is not volatility or the ratio of return to variation — it is the depth and duration of the fall it can sit through, which is why the drawdown-based measures deserve more weight than they usually get. A strategy with a Sharpe of 2.27 and a 7% maximum drawdown is one a household can hold; a Sharpe of 0.51 with an 18% fall may be better on paper and worse in practice, because the realised return on the second is set by whether the holder sold in the middle of it. That makes the drawdown the denominator that matters for any plan with a horizon, and it makes the choice between these strategies a question about the household rather than about the strategies. Two limitations are worth stating so that the numbers do not get over-trusted. The first is that all three measures are computed from a sample, so a three-year record produces a Sharpe ratio with a large standard error — enough that the ranking should be treated as a direction rather than a result. The second is that none of them says anything about the shape of the tail, and a strategy that has not yet experienced its bad event will report a wonderful Calmar right up to the point it does. That is the same fat-tail problem from R1 appearing in measurement rather than in modeling. The practical habit is therefore modest and useful: never compare two strategies or two funds on return alone; ask for the drawdown and, where available, the ratio of return to drawdown; and treat every one of these numbers as an estimate from a particular period rather than as a property. What the measures buy is not precision but a vocabulary for the trade-off — which is the last thing a household needs before writing the policy that turns all of it into rules. What each measure hides — Sharpe: punishes large gains as if they were losses · Sortino: breaks down with a zero denominator when there are no losing periods · Calmar: rests on a single observation — the worst one · All three: are estimates from a sample, and silent about the tail ← A three-year record gives a Sharpe ratio with an error large enough to reverse a ranking. Use these measures to structure a comparison, not to settle one — and always look at the drawdown beside the number.

How much evidence is a ratio?

A ratio of return to risk looks like a fact and behaves like an estimate, and its precision is much lower than the number of decimal places suggests. The standard error of a Sharpe ratio falls with the square root of the number of observations, which produces a rule that is easy to carry: the ratio’s **t-statistic** is approximately the ratio itself multiplied by the square root of the number of periods. At a yearly frequency, a Sharpe of 1.0 needs about four years to reach a t of 2, and a Sharpe of 0.5 needs about sixteen. At a monthly frequency the same arithmetic needs a hundred and forty-four months for the 0.5. The consequence is blunt — for most track records of the length a person will actually accumulate, the margin between a good Sharpe and a mediocre one is inside the error bars. The estimate is also biased upward by anything that smooths the return series, which is a property of the assets rather than of the statistic. Illiquid assets are valued by appraisal, so their reported returns are averages of surrounding periods, which suppresses measured volatility and therefore inflates the ratio. A private credit fund and a public equity fund with identical economics produce different Sharpe ratios, and the illiquid one looks better for a reason that has nothing to do with performance. The same effect appears in any strategy that holds stale marks, and it is one of the reasons an allocation can look beautifully diversified on paper until the reporting lags catch up. The practical discipline is to use these ratios for **ranking things measured the same way over the same period**, and never for making fine distinctions or for signing off on a strategy on the strength of the number alone. Two candidates measured over the same years with the same frequency can be compared. A single candidate with an attractive ratio and a two-year record has told you almost nothing, because it has not yet had the chance to produce the drawdown that defines its risk denominator — which is the specific failure mode that turns a Sharpe ratio from a summary into a sales document. Rough years of annual data needed for the ratio to be distinguishable from zero — Sharpe 1.0: ≈4 years for a t-statistic near 2 · Sharpe 0.5: ≈16 years — longer than most published records · Sharpe 2.0: ≈1 year, which is why short records with high ratios are the least trustworthy ← The most important line in any performance report is not the ratio but the length of the period it was computed over, and whether the series contains at least one episode of the kind that would test it.

The ratio’s hidden multipliers: annualising, autocorrelation and the search that produced it

The t-statistic rule in the previous read — the ratio multiplied by the square root of the number of periods — is only valid under an assumption that is quietly violated by most return series. The rule assumes the observations are **independent**, and monthly returns are not: momentum and valuation effects produce positive autocorrelation, and a fund holding illiquid assets marks them with a lag, which manufactures serial correlation whether or not the underlying values moved. When returns are positively autocorrelated, the effective number of independent observations is smaller than the count, so the t-statistic is overstated and the ratio looks more significant than the data supports. The practical consequence is the annualisation convention itself: multiplying a monthly Sharpe ratio by the square root of twelve is standard and is only correct under independence. With a serial correlation of, say, 0.2 — not unusual for a strategy holding less liquid positions — the correct annualisation factor is materially lower, and the reported ratio is inflated by something in the region of fifteen to twenty percent. That is enough to move a strategy from “promising” to “unremarkable”, and it is invisible unless the autocorrelation is checked. The second hidden multiplier is the **search**. Every strategy in this subject was arrived at by trying things: a parameter sweep, a filter, a universe change, a stop rule. If you tried twenty variations and report the best, the reported Sharpe is the maximum of twenty estimates, and the maximum of noisy estimates is biased upward even when none of the variations has any edge at all. The correction has a name — the deflated Sharpe ratio — and its content is simple: the benchmark a strategy must beat is not zero but the Sharpe that the best of N random attempts would have produced by luck, where that benchmark rises with the number of attempts and falls with the length of the sample. For a researcher who tried fifty configurations on five years of monthly data, the luck-benchmark is well above zero, which means that a reported Sharpe of 0.8 from such a search is close to what a coin would have delivered. The honest number to record alongside a result is therefore **the number of configurations tried**, which is exactly the discipline the trading-system lesson demands when it requires rules to be written before testing. The third is the difference between the ratio a strategy earns and the ratio its holder experiences. The Sharpe of a strategy is computed on a continuous full-size equity curve; the ratio a person realises depends on when they added, when they withdrew, and what they did during the drawdown — which is the behaviour gap the personal-finance lessons quantify. A strategy with a published Sharpe of 1.0 whose holders time their entries and exits badly can deliver a much lower figure to the household, and the paper number will still be quoted as though it were a property of the product. So the three multipliers to check before trusting a ratio are the autocorrelation, the count of attempts behind it, and whose cash flow it describes; a ratio that survives all three is comparable with another, and one that has not been checked is a number with a story attached rather than a measurement. Three multipliers on one Sharpe ratio — Reported: 1.0 annualised: Monthly ratio × √12, which assumes independent months · At 0.2 serial correlation: The correct annualisation is lower — the ratio is inflated by 15–20% ← · Best of 50 configurations: The luck-benchmark rises toward the reported number; record how many were tried ← · Strategy versus holder: The realised ratio depends on contributions, withdrawals and behaviour at the low The habit that makes all three checkable is a research log: the date each rule was written, the number of variations tested, and the autocorrelation of the return series. None of it is glamorous, and it is the difference between a backtest that reports a number and one that reports what the number means.

What you'll practise

A averages 12% with σ = 15.62%; B averages 10% with σ = 2.65%; the risk-free rate is 4%. What are the Sharpe ratios?

40 XP in the app · multi select

Sources

Practise this in the app →

Learn content is for education only — not individualized financial advice, a recommendation, or a solicitation to buy or sell any security. Options involve substantial risk. Examples are simplified and historical patterns never guarantee future results.