ClearViewLesson libraryWhat's new

Learn · Market Psychology · Thinking in Probabilities

Calibration and the Scenario Table

30 min read

A view becomes usable when it is written as scenarios with probabilities and a price implication, and confidence becomes trustworthy when it is scored — because a number you can be wrong about is what sizing, patience and hedges are actually built on.

From a view to a table

A scenario table has three rows and three columns, and writing it is the discipline. The rows are the cases: a bull case, a base case and a bear case, defined by the driver rather than by the price. The columns are the probability attached to each, the thing that decides it — a margin, a competitor, a refinancing date — and the implication for the position. The value of the exercise is not the arithmetic, which takes a minute; it is that the bear case has to be described rather than feared, and that the probabilities have to be written down where they can be scored. Two habits make the difference between a table that informs and one that decorates. The first is that the drivers have to be checkable. “Sentiment turns” is not a driver; “gross margin below 38% for two consecutive quarters” is, because a date and a number decide it. The second is that the probabilities should start from the base rate rather than from the story — the previous lesson’s discipline applied to the table — so that the bull case is not given a high weight because it is the most pleasant to imagine. A table where the three probabilities are 33% each is usually a table where nobody has done the work. The output is genuinely modest most of the time, and that is a feature. A position whose expected value computes to around 5% is one where the size should be small, the payoff comes from being right about the driver rather than from the arithmetic, and a single unexpected development can move every row at once. The table does not produce conviction; it produces a stated range, a stated downside, and a number that the position can be held against without the story being revisited every morning. Bull, base and bear — Bull: guidance beats, multiple holds — 25%: +40% · Base: in-line growth, multiple drifts — 50%: +8% · Bear: competitor wins, multiple compresses — 25%: −35% ← · Expected value: 0.25 × 40 + 0.50 × 8 + 0.25 × (−35) = +4.75% ← · What the number implies: A modest edge, a modest size, and the bear case priced in advance The bear column is the one that earns its keep. Writing −35% down, with the driver that would cause it, is the difference between a plan that can be checked and a hope that can be revised.

Keeping score on confidence

The second half of calibration is unglamorous and is where most of the improvement comes from: record forecasts with confidence levels, and check them. A trader who logs “beats estimates — 70% confident” and then counts how often the 70% forecasts were right discovers, usually, that they were right about 50% of the time. That is overconfidence made visible, and it is not a character finding: it is a measurement, and measurements can be improved by feedback in the same way a free-throw percentage can. The mechanism is worth being precise about, because it is why the exercise changes behaviour rather than merely describing it. Once the 70% forecasts are known to be right half the time, every subsequent number carries a correction, and the sizing that depends on a probability — the whole point of having one — stops being a fiction. A trader who believes a 55% view and a trader who believes an 80% view should hold different sizes; without scoring, both beliefs are uncalibrated, and the difference between them is a feeling rather than an input. The honest limits are two. Calibration takes volume, so a dozen forecasts say nothing and a few hundred say something, which is why the practice belongs alongside the trade journal rather than instead of it. And a well-calibrated forecast can still be a bad forecast: the world can be harder to predict than the confidence suggests, and a forecaster who is right 50% of the time on 70% calls is not merely badly calibrated, they may be taking on questions with too much noise in them. Both readings point at the same corrective, which is to narrow the range of what is forecast until the confidence level and the outcome agree. • Log forecasts with a confidence level and a date. • Score them in batches: calibration is a measurement, not a trait. • Correct every later number by the observed gap. • Feed the corrected probability into sizing, which is what the number is for. • Take volume into account: a dozen forecasts is noise, a few hundred is a calibration curve. The most common calibration error in investing is not overconfidence about the direction — it is overconfidence about the interval. People asked for a 90% range for next year’s index level produce ranges that miss reality far more than one time in ten, because the range was assembled from the plausible rather than from the distribution.

Scoring confidence, and the result that should humble you

If confidence is going to be kept score on, it has to be scored with a rule, and the standard one is simple: for each forecast, take the probability you assigned to what happened and square the distance from perfect certainty. A forecast made with ninety percent confidence that turns out wrong scores badly; one made with sixty percent confidence that turns out right scores mildly well. Averaging that score across many forecasts gives a single number that improves as your probabilities become more honest, and crucially it rewards being uncertain when uncertainty is warranted rather than rewarding confidence as such. Being right for the wrong reason does not raise your score much. The benchmark finding from the calibration literature is worth carrying into every scenario table, because it is remarkably stable across populations and decades. When people are asked for ranges they are ninety-eight percent sure contain the answer, the true value falls inside roughly **half to two-thirds** of the time. Overconfidence of that size is not a personality trait of traders; it is the base rate for humans setting intervals without a scoring rule. It is also not symmetric — most people’s ranges are too narrow rather than too wide, which is why the practical correction is to widen the range rather than to shift it. Applied to the bull-base-bear table in this lesson, that has two consequences. Your bear case is probably not bearish enough, and your probabilities are probably too concentrated on the middle. Both errors push sizing in the same direction, which is why the scenario table is not just a thinking aid: if you widen the range and thin the middle, the position that comes out is smaller than the one your confident base case wanted. That is the point at which calibration stops being an intellectual exercise and starts changing what you own. • Score confidence by the probability assigned to what actually happened, so honesty beats certainty. • Averaged over many forecasts, that score improves only as your probabilities become truthful. • People’s ninety-eight-percent ranges contain the truth roughly half to two-thirds of the time. • The usual error is ranges that are too narrow, not too wide — so widen the tails first. A calibration practice needs volume to be meaningful. Ten forecasts will not tell you anything; a hundred will show you whether your seventy-percent statements come true about seventy percent of the time, which is the only version of confidence that can be used for sizing.

Three overconfidences, and the one that costs

“Overconfidence” is used as though it were one thing, and the research separates three, with different causes and very different consequences. Naming which one is operating is what makes the fix possible, because the standard advice — be more humble — addresses only one of them and can make another worse. The first is **overestimation**: thinking you are better than you are, on average and in absolute terms. It is what people mean when they say a driver is overconfident. In markets it shows up as expected returns that are too high and as a belief in one’s own ability to beat the market, and the interventions are the obvious ones — a benchmark, an honest record, outside evidence. The second is **overplacement**: believing you are better than others rather than better than the truth. It is comparative, it is especially strong for tasks people consider easy, and it means the market can be full of people who each think they are above average and be right about the distribution while wrong about themselves. Its signature in trading is the willingness to take the other side of a trade on the grounds that the seller must not know what you know. The third is **overprecision**, and it is the one that does the most damage quietly. It is the belief that your estimates are more accurate than they are: your intervals are too narrow and your ranges too tight. The standard demonstration asks for a ninety percent confidence range around an estimate, and the truth falls inside it roughly half the time. Overprecision is not a belief about skill at all; it is a failure to represent uncertainty, and it is what turns a reasonable analysis into an oversized position, because the size of a bet is set by the width of the range rather than by the midpoint of the view. That is why the confidence rating the earlier read introduced does more work than it appears to. Recording a probability and then scoring it separates the three: a forecaster whose outcomes cluster near their stated confidence is calibrated, one whose outcomes land outside their stated ranges is overprecise, and one whose average stated probability is higher than their hit rate is overestimating. Each has a different remedy, and the remedy for overprecision is not humility but widening the intervals — which is a mechanical adjustment rather than an attitude. • Overestimation is about you against the truth; overplacement is you against others. • Overprecision is about the width of your estimate, not the height of your ability. • Ninety percent ranges contain the truth about half the time — the standard evidence. • Position size is set by the width of the range, so overprecision is what makes a bet too big. A single habit tests all three: attach a probability to each forecast and score it. The scoring does not require discipline or introspection, only arithmetic, which is what makes it more reliable than a promise to be more careful.

What you'll practise

Probabilities of 20% on a +50% case, 60% on a +5% case and 20% on a −30% case give what expected value?

40 XP in the app · multi select

Sources

Practise this in the app →

Learn content is for education only — not individualized financial advice, a recommendation, or a solicitation to buy or sell any security. Options involve substantial risk. Examples are simplified and historical patterns never guarantee future results.