Survivorship Bias
Drawing conclusions from what remains visible, when the thing you need to know is contained in what disappeared.
When not to use it
- Where the selection mechanism is genuinely unrelated to the outcome, which makes the sample usable and requires checking rather than assuming.
- As a way to dismiss any inconvenient finding, since every sample is filtered somehow and the question is whether the filter matters here.
- Where the missing population has been characterised and bounded, in which case the work has been done and the estimate stands.
Reach for something else instead
- Bounding — state a floor and a ceiling rather than a point estimate where the missing population is unknown.
- Prospective registration — record the population before outcomes are known, which removes the filter rather than correcting for it.
- -
Read more on the blog
- Nine cases, and none was fixed by a better modelTerritory 6 closes. Nine documented cases, three containing no AI at all, and not one where the remedy that worked was a more accurate system.
- AI incidents: two registers, 1,460 and 14,530The two main public AI incident databases count the same phenomenon and report 1,460 and 14,530. Neither is wrong. This is how to read an incident record, and what it cannot tell you.
- AI in hiring: 18 bias audits from 391 employersResearchers checked 391 New York employers against the world's first algorithmic bias audit law. Eighteen had posted an audit. Nearly every audit that existed reported passing.
- The pilot failed and the staff deployed it anywayEnterprise AI is measured by what organisations sanctioned. A separate literature measures what their employees actually use, and the two describe the same companies without ever being set against each other.
Further reading
- Wong et al. (2021), External Validation of a Widely Implemented Proprietary Sepsis Prediction Model — a failure mode that produces no artefact and is therefore invisible to any register.
- Pineau et al., reproducibility programme work, and Bouthillier et al. (2021), Accounting for Variance in Machine Learning Benchmarks — why a literature of reported results describes what was submitted.
Primary sources, listed so you can check the claims on this page rather than take them on trust.
Where people go wrong
- Treating a register of reported incidents as a count of incidents.
- Comparing surviving members of a population across time without noting that the population changed.
- Assuming absence of evidence in a filtered sample is evidence of absence.
At a glance
Where this sits
A starting point. Nothing needs to come before it.
Computed from the prerequisite graph, not assigned. How this works