Survivorship Bias
Drawing conclusions from what remains visible, when the thing you need to know is contained in what disappeared.
When not to use it
- Where the selection mechanism is genuinely unrelated to the outcome, which makes the sample usable and requires checking rather than assuming.
- As a way to dismiss any inconvenient finding, since every sample is filtered somehow and the question is whether the filter matters here.
- Where the missing population has been characterised and bounded, in which case the work has been done and the estimate stands.
Reach for something else instead
- Bounding — state a floor and a ceiling rather than a point estimate where the missing population is unknown.
- Prospective registration — record the population before outcomes are known, which removes the filter rather than correcting for it.
- -
Further reading
- Wong et al. (2021), External Validation of a Widely Implemented Proprietary Sepsis Prediction Model — a failure mode that produces no artefact and is therefore invisible to any register.
- Pineau et al., reproducibility programme work, and Bouthillier et al. (2021), Accounting for Variance in Machine Learning Benchmarks — why a literature of reported results describes what was submitted.
Primary sources, listed so you can check the claims on this page rather than take them on trust.
Where people go wrong
- Treating a register of reported incidents as a count of incidents.
- Comparing surviving members of a population across time without noting that the population changed.
- Assuming absence of evidence in a filtered sample is evidence of absence.
At a glance
Where this sits
A starting point. Nothing needs to come before it.
Computed from the prerequisite graph, not assigned. How this works