External Validation
An independent check of a model on data its developer did not choose, which is the only test that separates performance from the conditions it was reported under.
When not to use it
- As a one-time gate. A validation describes performance on one population at one time, and both change.
- Where the validating population is nothing like yours, in which case the result bounds the claim rather than confirming it.
- To dismiss a system outright. A poor external result raises the question of fitness for your setting; it does not settle it.
Reach for something else instead
- Local validation — evaluate on your own population before clinical or operational use, which answers the question that matters to you.
- Prospective evaluation — measure the system in live use against outcomes, rather than retrospectively against recorded ones.
- -
Read more on the blog
- The sepsis model caught 7% of what clinicians missedA sepsis warning system ran at hundreds of US hospitals before anyone outside the vendor validated it. The external check found the number that matters is not the one being reported.
- AI in medicine: 1,524 devices, 1.6% with trial dataThe FDA lists 1,524 AI-enabled medical devices. A review of 691 found 1.6% cited a randomised trial and under 1% reported patient outcomes. Medicare pays for about ten.
- 2.6 million robotic surgeries, none of them autonomousThe most-deployed medical robot makes no decisions. And the largest meta-analysis of its outcomes was co-authored by the company that sells it.
- 220 million miles, inside a boundary Waymo drewThe strongest safety evidence in physical autonomy, and the methodology that makes it honest is also what limits what it can tell you.
Further reading
- Wong et al. (2021), External Validation of a Widely Implemented Proprietary Sepsis Prediction Model — the canonical case, and the source of the 7% figure.
- Raji et al. (2021), AI and the Everything in the Whole Wide World Benchmark — why a developer-selected evaluation cannot license a general claim.
Primary sources, listed so you can check the claims on this page rather than take them on trust.
Where people go wrong
- Accepting a vendor-run study on customer data as external. The evaluators matter as much as the population.
- Reporting discrimination without calibration, which hides whether the stated probabilities mean anything.
- Treating deployment scale as evidence. Hundreds of installations tell you about procurement, not performance.
At a glance
Where this sits
A starting point. Nothing needs to come before it.
Computed from the prerequisite graph, not assigned. How this works