AtlasReasonSurvivorship bias, and the planes that never got counted

Bias · what the record leaves out · 7 min read

Survivorship bias, and the planes that never got counted

Abraham Wald's wartime armour study, the fund-return literature it foreshadowed, a simulation of vanishing failures, and three scenarios to try the rule on.

In short

Survivorship bias is drawing a conclusion from a group that has already been filtered by the very outcome you are trying to explain, without accounting for the cases that were filtered out. It is less a quirk of judgement than a fact about where data comes from: if failures leave no trace, whatever you measure in what remains will look better, safer or more skilled than the truth. The clearest early statement of the idea is Abraham Wald's Second World War work on where to armour bomber aircraft; the clearest modern demonstration is in the study of investment fund returns, where funds that close are quietly dropped from the record that later gets averaged.

An everyday example

A magazine profiles fifty highly successful founders and reports that most of them describe taking big, early risks. The advice that follows, take big risks early, sounds supported by fifty real people. It only looks that way because the founders who took the same big risks and failed are not in the sample; they are running a different business, or no business, and nobody profiles them under the heading of what worked.

The founders in the article are not lying and the magazine has not made anything up. The distortion is entirely in which group got measured: only the ones who survived long enough, and succeeded enough, to be worth interviewing.

The classic cases

During the Second World War, engineers studying where to add armour to bomber aircraft examined the planes that returned from missions and marked where the bullet holes clustered: mostly along the wings, the tail and the rear fuselage. The instinct was to reinforce those areas. Abraham Wald, working with the Statistical Research Group, argued the opposite: the returning planes were the survivors, and the holes on them marked the damage a plane could take and still make it home. The areas with few or no holes on returning aircraft, such as the engines and the cockpit, were the places a hit was likely to bring a plane down before it ever got counted. Wald's recommendation was to armour where the survivors were undamaged, not where they were damaged, because the untouched areas on a survivor were exactly the areas a non-survivor had probably been hit. Mangel and Samaniego's 1984 paper sets out the statistical method behind Wald's wartime memoranda in full, showing how the same reasoning gives an estimate of a plane's true vulnerability from data that only ever includes planes that returned.

Brown, Goetzmann, Ibbotson and Ross's 1992 study applies the identical logic to a very different record: databases of investment fund returns. A fund that performs badly for long enough typically closes, merges into another fund, or simply stops reporting, and once that happens it is often removed from the database that is later used to calculate average historical performance. A database queried today, built from funds still operating today, is therefore not a record of how funds as a group performed; it is a record of how the funds that happened to survive performed, with the weakest performers filtered out along the way precisely because they performed weakly. The paper's conclusion, in its own title, is that this survivorship problem is a real and measurable distortion in how fund performance gets studied, not a theoretical worry.

A sampling problem before it is a thinking problem

It helps to separate two things that survivorship bias is often blamed for at once. The first is a fact about data: any record built from survivors is missing the non-survivors, and that is true whether or not any person ever reasons about it. The second is what a person then concludes from that incomplete record, which is where the psychology comes in: reading a survivors-only record as though it were the whole story, and drawing lessons from it accordingly. The founders' magazine and the fund database both start from the same structural gap; the mistake happens at the point a reader treats what is left as though nothing were missing.

Does it replicate?

Replication grade: Strong: a structural sampling effect, documented across many independent datasets

Survivorship bias is not tested for replication the way a laboratory result is, because it is not really a claim about how people's minds work under uncertainty; it is closer to a mathematical necessity. Whenever a selection step removes cases based on the very outcome later being measured, and the removed cases are not added back into the analysis, the remaining average will be biased in the direction of the outcome that caused the removal. That has been documented again and again since Brown, Goetzmann, Ibbotson and Ross's 1992 paper, in hedge fund and mutual fund research, in studies of business survival, and in any record, financial or otherwise, that only lists what is still around to be listed.

What is worth debating case by case is the size of the distortion, which depends entirely on how the record was built: how funds or cases drop out, how completely, and whether anyone corrected for it afterward. That is a question about a specific dataset, not about whether the underlying logic holds, which it reliably does.

A simulation of vanishing failures

A simulation of survivorship bias in invented fund returnsSimulated, not real, data: of 60 invented funds with random returns, 2 post a run poor enough to close and drop out of later listings. The average return of every fund that ever launched is 3.4%; the average of only the 58 still reporting is 4.0%, higher only because the worst performers are no longer in the sum.0%3.4%Every fund that launchedn = 604.0%Only funds still reportingn = 58Average return (simulated)
Simulated, not real, data: of 60 invented funds with random returns, 2 post a run poor enough to close and drop out of later listings. The average return of every fund that ever launched is 3.4%; the average of only the 58 still reporting is 4.0%, higher only because the worst performers are no longer in the sum.

The diagram on this page is a simulation, not real fund data. It invents a set of funds with random returns and lets the worst of them close and drop out of a later listing, then compares the average return of every fund that ever launched with the average of only the funds still reporting. No skill or insight is added anywhere in the simulation; the only thing that changes the second number is which funds are still there to be counted.

Try it: three scenarios

Three short situations built around the same shape: a group has already been filtered by an outcome, and a conclusion is drawn only from what is left.

1. A magazine studies the daily habits of 50 highly successful founders and finds that most wake before 6am. What is missing from this study before it can support the advice "wake up early to succeed"?
Why

Without knowing how common early waking is among people who tried and did not succeed, and among the general population, an early-waking habit among survivors cannot be told apart from an early-waking habit that has nothing to do with success.

2. A hospital finds that patients who received a certain surgery and lived to their one-year check-up report high satisfaction. Why is this alone not evidence the surgery is safe?
Why

Anyone who died or was too unwell to attend the check-up is absent from the satisfaction count, which is drawn only from the survivors: exactly the group least likely to report a bad outcome.

3. A city looks at its oldest buildings, still standing after 150 years, and concludes that old construction methods were sturdier than modern ones. What does this argument leave out?
Why

The surviving old buildings are, by definition, the ones sturdy or lucky enough to still be standing; the ones that fell down or were torn down long ago left no building to inspect, so they are silently absent from any conclusion drawn from what remains.

Nothing you type on this page leaves your browser tab: no answer, guess or score is sent anywhere or stored.

How to catch it

The question to ask is always the same: what happened to the cases that are not in front of me?

  • Ask how the group you are looking at got assembled, and specifically what would have had to happen for a case to be missing from it.
  • For any record of success, ask for the matching record of attempts, not only of successes: how many tried the same thing and are not being counted.
  • Be suspicious of any database, leaderboard or list that only shows current or still-operating entries, with no note of what used to be on it and left.
  • When a group is described as "the ones still standing," "still in business" or "still with us," ask what standing, in business or with us had to mean for something not to make the list.
  • Remember that the missing cases are usually the ones that would tell against the conclusion; that is exactly why they are the ones missing.

Check yourself: three questions

1. In the wartime aircraft study, why did Wald recommend armouring the areas that were LESS damaged on returning planes?
Why

The returning planes were survivors. Few holes in an area on survivors did not mean that area was safe; it more likely meant a hit there tended to bring a plane down before it could return and be counted.

2. What does Brown, Goetzmann, Ibbotson and Ross's 1992 paper say about databases of fund performance built from funds still operating today?
Why

The paper's own title names the problem: a database drawn from surviving funds is missing the funds that closed, typically the weaker performers, which inflates the average return the database appears to show.

3. Why is survivorship bias described on this page as closer to a mathematical necessity than to a debated psychological finding?
Why

The distortion follows from how the record was built, not from anyone's reasoning; given that structure, the bias in the remaining average is a predictable consequence rather than something that sometimes shows up and sometimes does not.

Sources

Text on this page is original to MyTestAtlas, written from the studies listed. The diagram is drawn by this site and is not a copy of any published figure.