Learn · Market Psychology · Thinking in Probabilities
Survivorship and the Sample You Actually See
Every sample available to a trader is the subset that survived, and selection is not random — so the first question about any record is what was excluded, and the second is how many attempts it was chosen from.
Why the missing data is not random
The cleanest demonstration is not financial. During the Second World War, statistician Abraham Wald was asked where to add armour to bombers returning from missions, and the aircraft came back with more damage in some areas than others. The obvious reading was to reinforce the most-damaged regions. Wald’s answer was the opposite: the sample consisted only of planes that returned, so the regions with little damage were the regions where a hit was fatal. The missing aircraft were the data, and they were missing precisely because of the variable being studied. Fund databases have the same structure. A fund that closes disappears from the record, and funds close after performing badly — so the average of the funds still listed is upward biased by construction. The bias is measurable and has been for decades: studies comparing complete samples with survivor-only samples find the excess return overstated by roughly half a point to a point and a half a year for mutual funds, and considerably more in hedge fund databases, where a second bias compounds it. Many funds report to a database only after a period of good performance, so the history that appears is selected at the moment of entry as well as at the moment of exit. The third form is the one that decides whether a backtest means anything, and it is a property of how the strategy was found rather than of the data. Test enough variations and some will look excellent by chance. A hundred independent ideas evaluated at a 5% significance level will produce about five that look significant when none is, and the one that gets shown is the best of the hundred — so the reported result is a maximum, not a sample. The remedy is to know the number of attempts that produced the result, and to raise the evidence bar with it: the same literature that documents this recommends a threshold closer to a t-statistic of 3 rather than the conventional 2 for the factors that survive such searching. One directory, two averages — Funds that began: 500 · Closed or merged: 300, averaging 3.2% a year ← · Still listed: 200, averaging 8.0% a year · Published average: 8.0% · Average across all 500: 5.12% ← · Overstatement over 25 years on $10,000: about $33,600 Both averages are computed correctly. The difference is the denominator: the directory reports the funds that are there to be reported, and the returns of the ones that are gone have to be recovered from somewhere else or estimated.
Reading your own record the same way
The bias is not only a problem with other people’s statistics. A trader’s own history has a denominator too, and it is usually invisible. Consider what a journal actually contains: if it was started after a good stretch, everything before it is missing; if losing trades were not logged at the time because the trader was avoiding the screen, the record is a sample of the trades the trader was willing to write down; and the trades remembered vividly are the ones with a strong feeling attached, which is availability rather than frequency. The audit from the mastery rung works only on a record with a complete denominator, which is why the journal has to be written before the outcome rather than after it. Three questions resolve most of the cases a trader will meet. What is missing from this sample, and why did it leave? For a fund average, that is closed funds; for a backtest, the failed variants; for a personal record, the unlogged trades. How was the sample selected — was it everything, or the part that reached a threshold, or the part the researcher could get? And how many attempts produced the result being shown, because a record selected from fifty trials is a different object from one selected from one. None of this requires a statistical test; it requires looking for the denominator before looking at the number. The practical consequence is a habit of mind rather than a calculation: prefer evidence whose selection is transparent to evidence that is impressive, treat the absence of an excluded set as a warning rather than a reassurance, and set the bar higher for anything found by searching — including your own backtests, which are the easiest place to hide the number of attempts from yourself. The base-rate lesson from earlier in this rung supplies the same instruction from the other side: classify the situation, ask how the class resolved, and only then read the story that is being told about the instance in front of you. • Ask what left the sample and why — closed funds, failed variants, unlogged trades. • Ask how the sample was selected, and whether the selection threshold is the variable under study. • Ask how many attempts produced the result; a maximum of many trials is not a sample of one. • Raise the bar with the number of attempts, towards a t-statistic of 3 for searched effects. • Prefer transparent selection to impressive numbers, and treat a missing denominator as a warning. The mirror error is assuming that every good record is built on a hidden denominator. Sometimes the sample is complete and the result stands. The discipline is to ask the question rather than to answer it in advance — the same stance the audit takes toward a process that turns out to be worth nothing.
The same bias lives in published research
Survivorship is easiest to see in a directory of funds and hardest to see in a journal article, and the direction of the distortion is identical. A finding gets published when it is statistically interesting, and a result that fails to reach significance is filed away; so the body of published findings is a sample selected on the very property being tested. Do that across thousands of independent researchers testing thousands of combinations, and a certain number of false positives is guaranteed — not because anyone cheated, but because “significant at the five percent level” is defined as the rate at which noise is expected to look real. In finance this shows up as the **factor zoo**: hundreds of published characteristics each shown to predict returns, many of which fail to hold up out of sample. The statistical response is now standard practice and easy to carry over to your own reading. When many hypotheses are tested against the same data, the threshold for significance has to be raised — the work on this argues that a t-statistic around three, rather than the conventional two, is the more honest bar for a newly claimed effect in asset pricing. The practical version of that rule for a reader is simple: a single study, testing many things, reporting one strong result, has done something that needs replication before it needs belief. Replication by an independent team on a different sample is worth more than a higher t-statistic from the one sample — the same logic that makes the case studies in this curriculum useful only as illustrations rather than as evidence. There is a second selection effect specific to fund and strategy data that is worth knowing by name. **Backfill bias** describes databases that add a manager once it has a track record worth reporting, meaning the early, unimpressive years are never recorded, and the index the manager is measured against contains only managers who lasted long enough to be included. Together with survivorship it can lift a reported average by an amount comparable to the entire supposed skill being demonstrated. Any aggregate performance figure — an average hedge fund return, a peer-group median, a strategy index — should be read with the same question the fund directory raises: how did the members of this sample get into it? Three questions to ask of any performance statistic: who is in the sample, how did they get in, and how many attempts were made that did not end up in the sample. The third is the one that turns a striking result into an ordinary one.
The databases are survivors too: backfill, delisting and the missing funds
The previous read showed that any sample assembled by whoever survived is biased, and the same logic reaches every dataset a learner is likely to consult. Take a commercial database of hedge funds, which is the standard source for the claim that a category of managers earned some rate. Funds report voluntarily, and they tend to begin reporting after a good run, which means the early history is filled in after the fact — the industry calls it **backfill bias**, and its measured effect on reported returns has been estimated in the region of several percentage points a year. The same funds tend to stop reporting when performance sours, and the database keeps the last good number rather than recording a final collapse — **delisting bias**, worth roughly the same order of magnitude in the opposite direction of honesty. The combined effect is that a simple average across the surviving records overstates what an investor would have earned, not because any number is fake but because the sample was assembled by the outcome. The same three selection effects operate inside individual datasets rather than across them. An index of companies has a membership rule, and membership changes: a benchmark of large-cap names kept the successful and dropped the failures, so its long history describes a set of companies chosen partly for having done well. A public dataset of strategies reported by their authors contains the ones that were deemed worth publishing. And a set of academic findings about a market anomaly is a set of results that survived the review process, which is the publication bias the previous read named. The practical test in every case is the same question, asked in the same form: **what had to be true for this observation to appear in my sample, and would an observation like it have been recorded if it had gone the other way?** A fund that closed is not in the sample; a study that found nothing is not in the journal; a trade that you read about because it worked is not accompanied by the fifty that did not. Two habits follow, and both are cheap. The first is to have a preference for datasets with a **stated and enforced inclusion rule** — a survivorship-free index, a database that retains delisted records with a final return, a pre-registered study — because the rule is what tells you what is missing. The second is to treat any surviving entity’s history as an **upper bound** rather than an estimate: the surviving funds overstate the category, the surviving companies overstate the strategy, and the surviving trades in your own memory overstate your process. That is not pessimism; it is knowing which direction the error runs. When the bias has a sign, the honest estimate is a range with the recorded number at one end, and the interval is what should be used when it feeds into a decision — which is the same conclusion the sample-size arithmetic produced for a base rate, reached through the entry conditions rather than through the variance. • Backfill bias: funds start reporting after a good run, and the early history is filled in later. • Delisting bias: funds stop reporting when performance sours, and the last good number survives. • Index membership rules mean a benchmark’s history was partly chosen by its outcome. • Ask of any observation: what had to be true for it to be in my sample at all? • Prefer datasets with an enforced inclusion rule, and read surviving histories as upper bounds. The link to the audit lesson is the useful one: your own journal is a dataset with a selection rule too, and the rule is whatever you chose to write down. Recording the declined setups and the abandoned ideas is how a personal record becomes survivorship-free, which is the same fix the database vendors apply to theirs.
What you'll practise
200 surviving funds averaged 8.0% and 300 closed funds averaged 3.2%. What was the average across all 500?
40 XP in the app · multi select
Sources
- Survivorship bias in performance studiesBrown, Goetzmann, Ibbotson & Ross (1992), Review of Financial Studies
- Multiple testing and the cross-section of expected returnsHarvey, Liu & Zhu (2016), Review of Financial Studies
- A method for estimating plane vulnerability based on damage of survivorsWald (1943); Mangel & Samaniego (1984), “Abraham Wald’s Work on Aircraft Survivability”
Learn content is for education only — not individualized financial advice, a recommendation, or a solicitation to buy or sell any security. Options involve substantial risk. Examples are simplified and historical patterns never guarantee future results.