Unrecorded Stratifier
A variable that plausibly changes a result, known to be unevenly sampled, cheap to record, and absent from the record, which makes an aggregate uninterpretable rather than merely imprecise.
When not to use it
- Where the stratifier was recorded and reported, which is the condition the concept checks for.
- As an argument for exhaustive stratification, which produces underpowered subgroup analyses that mislead in their own way.
- Where no plausible mechanism connects the variable to the result, in which case its absence is ordinary rather than load-bearing.
Reach for something else instead
- Stratified reporting as a publication condition — cheap, prospective, and does nothing for existing literature.
- Purpose-built stratified benchmarks — the only remedy that works on models already trained and deployed.
- -
Read more on the blog
- The sepsis model caught 7% of what clinicians missedA sepsis warning system ran at hundreds of US hospitals before anyone outside the vendor validated it. The external check found the number that matters is not the one being reported.
- AI in medicine: 1,524 devices, 1.6% with trial dataThe FDA lists 1,524 AI-enabled medical devices. A review of 691 found 1.6% cited a randomised trial and under 1% reported patient outcomes. Medicare pays for about ten.
- 232 studies, and 1.3% recorded skin typeA systematic review found AI detecting skin cancer at 90% accuracy across 232 studies. Almost none of those studies recorded who the patients were, and the ones that checked found performance dropping to chance.
- Good evidence, and eight different failuresTerritory 11 closes. This was the field chosen because its evidence is strong, and every subject produced a failure that stronger evidence would not have prevented.
Further reading
- Wong et al. (2021), External Validation of a Widely Implemented Proprietary Sepsis Prediction Model — an independent evaluation on the population that mattered, following internal validation that did not break it out.
- Daneshjou et al. (2022), Disparities in dermatology AI performance on a diverse curated clinical image set — a purpose-built stratified benchmark, which is the only retrospective remedy available.
Primary sources, listed so you can check the claims on this page rather than take them on trust.
Where people go wrong
- Treating an uncharacterised aggregate as a figure with wide error bars rather than as a figure with an unknown referent.
- Assuming an imbalance averages out, when the studies were drawn from the same skewed sources.
- Reading the absence as neutral when every study that measured found an effect in the same direction.
At a glance
Where this sits
A starting point. Nothing needs to come before it.
Computed from the prerequisite graph, not assigned. How this works