ClearViewLesson libraryWhat's new

Learn · Market Psychology · Behavioral Biases

Hindsight, Outcome Bias and the Verdict on a Decision

30 min read

A decision is judged on what was knowable when it was made; the outcome is a draw from a distribution, so scoring the two together means learning the wrong lesson from both the wins and the losses.

Hindsight: the past is not as readable as it looks

Hindsight bias is the tendency, once an outcome is known, to believe it was more predictable than it was. It was demonstrated cleanly long before it had a market application: give people a description of a historical event and the outcome, and they rate the outcome as having been likely; give the same description with a different outcome and they rate that one as likely instead. The event does not change; the judged probability does. The information that was actually available at the time — the conflicting reports, the ambiguity, the case that could have gone the other way — is quietly deleted from the reconstruction. In markets the effect wears a familiar costume: “the bubble was obvious”, “the crash was written on the wall”, “anyone could see the company was a fraud”. Each claim is made after the fact and each is unfalsifiable, because the evidence for it is drawn from the outcome it is meant to explain. What the claim costs is not accuracy about the past but calibration about the future: a trader who believes the last move was predictable has no reason to record probabilities, and a record of probabilities is the only thing that can show them whether they are any good at this. There is a second cost that is easy to miss. Hindsight does not merely make the past look easy; it makes the present look legible too. If crashes are always obvious and recoveries are always obvious, then the correct action is always available and the only question is courage, which is a story that ends in oversized positions taken into genuinely ambiguous situations. The honest version is that the market is usually ambiguous while it is happening and obvious while it is being described. The same fact, two verdicts — Event: an index falls 12% over three weeks: What was knowable: mixed data, a rise in credit spreads, two conflicting earnings reports ← · Judged in the moment: A probability statement: perhaps 30% this quarter · Judged a month later: “It was inevitable” — a probability near 100% ← · What changed: The outcome, and nothing else about the information The test is mechanical rather than philosophical: write down the estimate, with a number, before the event. A trader with a stack of written 30% forecasts that came true three times in ten is calibrated; a trader who “saw it coming” has a memory.

Outcome bias: grading the decision by the draw

Outcome bias is the tendency to judge the quality of a decision by the quality of its result. Its classic demonstration is simple: describe a decision made by a doctor or an investor, then vary only the outcome, and the quality rating of the decision moves with the outcome. The decision was identical and the rating changes, which means the rating was never about the decision. In a domain where outcomes are noisy this is not a minor error — it is the error that misdirects the entire learning loop. The cost runs in both directions. A good process that produces a loss is punished: the trader changes the rule that worked, or abandons the system, at precisely the moment the system is behaving as designed. A bad process that produces a win is rewarded: no stop, doubled size, a lucky entry, and the record now says the shortcut works. Because the lesson taken is the opposite of the true one in both cases, the trader who evaluates by results gets systematically worse over time while feeling increasingly certain about what they have learned. The correction is a separation of the two axes. A decision is good or bad given what was knowable at the moment it was made and the distribution of outcomes it implied; the result is a draw from that distribution, and a single draw carries very little information about it. So the review asks two questions in order. Was the process followed — was the setup written, was the size from the plan, was the stop where the plan said? Only then: what does the result tell us, in a sample large enough to have a signal? Fifteen trades is the start of a sample; one is an anecdote, and treating an anecdote as a verdict is how the loop breaks. • Freeze the information set: judge the trade on what was knowable at entry, not on what followed. • Score the process first — setup, size, stop, exit rule — and the result second. • Keep a written estimate before the event, because hindsight deletes the ambiguity afterwards. • Only read results in samples; one trade is a draw from a distribution. A long run of good results is not evidence that the process is good, because the market is the largest contributor to most early records. The mirror case is worse in practice: a process that is working will produce losing months, and evaluating it monthly guarantees abandoning it in the month it did nothing wrong.

Reviewing a decision without rewriting it

The practical difficulty with separating decision quality from outcome is that the review happens after the outcome is known, by the same mind that made the decision. The fix is procedural and it works because it moves the information: **write the review criteria before the trade is opened.** If the entry, invalidation and size are recorded at the start, then the review is a comparison rather than a reconstruction, and the question becomes whether the plan was followed rather than whether the trade was a good idea. It is the difference between an audit and a memoir, and only one of them can be wrong. The second half of a fair review is to grade against a **sample of outcomes rather than the one you got**. A decision made with a forty-percent win rate is supposed to lose six times in ten; there is nothing to correct in the process when it loses, and there is nothing to celebrate when it wins. The way to keep that straight is to ask what the same decision would have produced over a hundred repetitions — a good bet that lost is a good bet, and a bad bet that won is a bad bet that has not been paid for yet. This is not a licence to ignore losses. It is a rule for deciding which losses are information: losses that came with the plan violated are evidence, and losses that came with the plan followed are the cost of doing business, priced in advance. The failure mode on the other side deserves equal attention, because it is where most disciplined traders damage themselves. A good process that produces a bad outcome is frequently “fixed” — a parameter tightened, a filter added — and each fix narrows the system toward the trades that already happened. The test for whether a change is legitimate is whether the same change would have been identifiable from the record *before* this trade, or whether it exists because of this trade. A pre-committed review cadence, on a schedule rather than after every loss, is what makes that test possible: the sample gets assessed as a sample, and the modification is made once, deliberately, on evidence rather than in reaction. The one-sentence version: record the plan at entry, review on a schedule rather than on a reaction, and ask what a hundred repetitions of the same decision would have produced before deciding the process was wrong.

The predictions ledger: scoring the forecast, not the trade

The pre-trade record in the previous read protects a decision from being rewritten after the fact. A **predictions ledger** does the same job for the beliefs underneath it, and it is the only instrument that can answer the question the probability lesson raised: whether your sense of what is likely has any content. The format is deliberately austere. Before acting, write the claim in a form that can be checked — not “I am bullish on the index” but “the index closes higher than today in three months, probability 65%” — with the date, the probability, and the information it depends on. The point of stating a probability rather than a direction is that it can be scored: a set of claims recorded at 70% should come true about seven times in ten, and the deviation is the measurement. Without the number, every outcome can be absorbed as roughly what was expected, which is how a decade of trading can leave someone with no idea whether their forecasts are informative. Scoring needs a rule fixed in advance, because the scoring itself is a place where hindsight operates. Three rules make the ledger hard to game. The claim is resolved on the date written, not when it looks decided; the resolution is mechanical, so that no judgement is exercised at the moment of grading; and the whole set is scored, including the claims that turned out to be uninteresting, because the ledger’s value depends on the denominator. A ledger scored selectively is a memory with extra steps. What the honest version produces is a **calibration curve**: group the claims by the probability attached and count how often each group came true, and the diagonal is the target. Miscalibration is extremely common and it is directional: probabilities near the high end tend to be overstated, with 90% claims landing around 70% to 80%, and the underconfidence at the low end is equally common. That asymmetry is the same overprecision the overconfidence lesson measures with interval questions, showing up in the forecasts a trader actually makes. The reason the ledger matters more than the trade journal is that it separates two failures that look identical in a set of results. A trader can be well calibrated and lose money through poor risk management, or badly calibrated and make money for a while on a market that happened to rise. The trade journal shows the P&L; only the ledger shows whether the forecasts had content, and the two diagnostic questions are different: the journal asks whether the process was followed, and the ledger asks whether the beliefs were any better than the base rate. Run together, they distinguish skill from the current regime. And there is a practical by-product that makes the exercise worth its friction: writing a probability before acting forces the estimate to be explicit, and an explicit estimate can be compared with the price, which is the entire mechanism by which a view becomes a position rather than a feeling. The ledger, and what the diagonal means — A claim that can be scored: “Higher in three months, probability 65%” — dated, with the information it rests on · Resolve on the date written: Mechanically, so no judgement is exercised at grading time · Score the whole set: Including the ones that became uninteresting — the denominator is the point ← · Read the calibration: 90% claims landing at 70–80% is the usual pattern, and it is directional ← The link to the base-rate lesson is the reason to bother: a ledger converts “I am usually right about this sort of thing” into a measured hit rate, which can then be conditioned on the setup exactly as a trade base rate is. It is also the only way to find out whether the confidence that sizes a position is calibrated or merely felt.

What you'll practise

A rule-based system with positive expectancy loses fifteen trades in a row out of sixty. What is the correct reading?

30 XP in the app · multi select

Sources

Practise this in the app →

Learn content is for education only — not individualized financial advice, a recommendation, or a solicitation to buy or sell any security. Options involve substantial risk. Examples are simplified and historical patterns never guarantee future results.