ClearViewLesson libraryWhat's new

Learn · Risk & Sizing · The Tail and the Cost

The Threshold and the Tail

35 min read

Value at risk is a quantile — the loss not exceeded 95% of the time — so it says where the tail starts and nothing about how bad it gets. Expected shortfall averages what lies beyond it: on a five-scenario distribution the 95% threshold is a 40% loss and the shortfall is 44%, or $440,000 on a million. Because tails are fat and probabilities are estimates, the useful exercise is a written stress test with named scenarios, a total loss, and a pre-committed response.

A quantile describes the edge of the tail, not the tail

Value at risk answers a narrow question: what loss is not exceeded at a given confidence level? At 95% on the distribution above, that is a 40% loss — the worst 5% of scenarios. It is a useful number and it is systematically misunderstood, because it is a **quantile**, which means it describes where the tail begins and says nothing about its depth. Two portfolios can share a 95% value at risk of 40% while one of them contains a scenario twice as bad as the other, and the measure cannot tell them apart. **Expected shortfall** — also called conditional value at risk — fixes the omission by averaging the losses beyond the threshold. On the same distribution it is 44%, or $440,000, against a threshold of 40% and $400,000. The relationship between the two numbers is itself informative: shortfall is always at least as large as the threshold, the gap between them says how heavy the far tail is, and a large gap is a warning that the scenarios beyond the threshold are much worse than the threshold suggests. That is the case in almost every market distribution, which is why regulators moved from value at risk to expected shortfall in the Basel framework. Both measures share a deeper problem, which is that they are computed from estimated probabilities. The probabilities in the table above are a judgement, and they are almost certainly too low for the extreme scenarios, because the sample period used to estimate them did not contain a once-a-generation event. Fat tails, in the R1 sense, mean the model is worst exactly at the point being modelled — so the honest use of these numbers is as a discipline for comparing portfolios and for forcing a conversation about the worst case, not as a measurement of what can happen. Threshold against tail — Value at risk, 95%: 40% — $400,000 · Expected shortfall beyond it: 44% — $440,000 ← · The gap between them: 4 points, and it measures how heavy the far tail is · Expected loss across all scenarios: 11.55% — $115,500 a year A distribution where the shortfall is only slightly above the threshold is one with a thin tail. A large gap means the scenarios beyond the threshold are much worse than it, which is the case for equities.

The stress test is a written exercise

Because the numbers are estimates, the useful version of this lesson is procedural rather than statistical. A **stress test** is a written exercise with four steps that take an hour rather than a model: name a small set of specific scenarios, estimate what each does to the portfolio holding by holding, add the losses to get a total, and compare that total with what the household could actually absorb. The scenarios should be historical rather than hypothetical — 2008, March 2020, 2022, the 2010 flash crash, a 2018-style volatility spike — because specific past events carry a mechanism with them and force the estimate to be about positions rather than about feelings. The step that people skip is the one that makes it a risk exercise rather than a spreadsheet: comparing the total with **capacity**. A 40% loss on a portfolio that funds a house purchase in eighteen months is not a risk to be tolerated but a plan to be changed, and the whole point of running the scenario is to discover that before it is a fact. Capacity includes the buffer, the guaranteed income, the time until the money is needed and the household tolerance — which is why this step cannot be delegated to a formula. Then the response has to be written down in advance. If a 25% loss would trigger a change — trimming the concentrated position, halting new risk, adding a hedge — that action belongs in the document with the trigger, not in the decision-making of a frightened holder at the bottom. The final property of a good stress test is that it is repeated on a schedule, because the portfolio changes, the exposures drift, and a document from three years ago is a description of a different portfolio. The four steps — 1. Name the scenarios: 2008, March 2020, 2022, the flash crash — with mechanisms attached · 2. Estimate each holding: position by position, at the scenario, not at the index · 3. Add to a total: a number, and it will be worse than the sum of the obvious ones ← · 4. Compare with capacity and pre-commit: buffer, horizon, guaranteed income — then the response, in writing A stress test whose scenarios are all plausible is not a stress test. The scenarios worth writing down are the ones that would change the plan, because those are the ones the plan is supposed to survive.

The three settings, and one property VaR lacks

A value-at-risk figure is meaningless until three settings are named alongside it, and reports that omit them are comparing different things. The **confidence level** decides where in the distribution you are describing: a 95% figure is the loss exceeded one day in twenty, and a 99% figure is the loss exceeded one day in a hundred, and the second is a much larger number. The **horizon** decides the window: a one-day and a ten-day figure differ by roughly the square root of ten if returns are independent, which is the standard scaling and is also a deliberate understatement, because losses cluster rather than arriving independently. And the **method** decides what is being measured at all — a historical simulation uses the past distribution, a variance-covariance approach assumes a shape, and a Monte Carlo approach assumes a model of the dependencies. The third setting is where most failures of the measure live. Scaling a one-day number by the square root of time assumes that consecutive days are independent, and in the states that produce breaches they are not: volatility clusters, and a day bad enough to reach the threshold is usually followed by more like it. So the ten-day figure built by scaling is systematically too small in the tail — exactly the same criticism this lesson makes of the threshold itself, arriving from a different direction. There is also a formal property that value at risk has and expected shortfall does not, and the direction surprises people. Value at risk is **not subadditive**: there are portfolios where the value at risk of the combined position is greater than the sum of the parts, which means the measure can report that diversification increased risk. A quantile depends on where the mass sits, so combining two positions can move the quantile without any real increase in danger. Expected shortfall — the average of the losses beyond the threshold — is subadditive and coherent, which is the technical reason regulators moved toward it and the practical reason a single VaR number should never be the only tail measure in a report. • Name the confidence level, the horizon and the method before quoting any number. • Square-root-of-time scaling understates long-horizon risk because losses cluster. • Value at risk is not subadditive — it can report that combining positions increased risk. • Expected shortfall is coherent; use both, and treat the difference as information.

How many breaches should there be?

A risk number is a forecast, and a forecast can be checked. A one-day figure at a ninety-five percent confidence level is a claim that the loss will exceed it on about one day in twenty — roughly thirteen days in a trading year. That count is observable, which is what turns a risk model from an assertion into something that can be monitored, and monitoring it is how a model is retired before it is retired for you. The simplest backtest is a tally. Record each day whether the loss breached the predicted figure, and compare the count with what the confidence level implies. The uncertainty around that count is not small: over two hundred and fifty days the expectation is about thirteen breaches, and a plausible sampling range runs from the high single digits to the high teens. A count of sixteen is not evidence that the model is broken; a count of thirty is. The count is the crudest of three tests, and the other two are more informative. The second is **clustering**: breaches should be spread out, so a model that produces them in runs has missed a change in the volatility regime rather than merely underestimated its level. Three consecutive breaches matter more than the same three spread across a quarter, and the check is simply whether a breached day is more likely to be followed by another than chance implies. The third is **size**: given that a breach happened, how far past the figure did the loss go? A model whose breaches are typically a little over the line behaves differently from one whose breaches are multiples of it, and it is the second kind that ends accounts. Regulatory practice offers a design worth borrowing, whatever the scale. Supervisors have classified breach counts into zones — a green region consistent with the model, an amber region that warrants attention, and a red region implying the model is inadequate, with consequences attached to each. The transferable principle is not the thresholds but the structure: a count, a presumption about what a breach means, and a response decided before the count arrives rather than after. The honest limit is the one that applies to every statistical check in this subject. A model can pass every backtest it is given and still fail, because the sample that tested it did not contain the event; that is not an argument against testing, it is the reason the stress test sits beside the statistical measure rather than being replaced by it. A red-zone count is a reason to reduce exposure, not a reason to re-tune the model until the count looks better — re-tuning is how the model is fitted to the past it was supposed to describe. The practical form for a private account is two columns in a log: the predicted loss and the actual one. The count, the clustering and the average overshoot are then available for free, and the question of whether the risk estimate is worth anything has an answer that is not a matter of opinion. • A ninety-five percent daily figure implies roughly thirteen breaches a year — countable. • Clustered breaches point to a regime the model missed; isolated ones may be sampling. • The average size of an overshoot says more about tail risk than the count does. • Set the response to a red-zone count in advance, and do not re-tune the model to fix it. Two columns and a monthly count cost nothing and answer the only question that matters about a risk estimate: has it been describing the losses, or has it been describing a calmer market than the one that arrived?

What you'll practise

Five scenarios: 4% at p = 0.50, 12% at 0.30, 25% at 0.15, 40% at 0.04, 60% at 0.01. On $1,000,000, what is the 95% value at risk and the expected shortfall?

40 XP in the app · multi select

Sources

Practise this in the app →

Learn content is for education only — not individualized financial advice, a recommendation, or a solicitation to buy or sell any security. Options involve substantial risk. Examples are simplified and historical patterns never guarantee future results.