External Validation
An independent check of a model on data its developer did not choose, which is the only test that separates performance from the conditions it was reported under.
When not to use it
- As a one-time gate. A validation describes performance on one population at one time, and both change.
- Where the validating population is nothing like yours, in which case the result bounds the claim rather than confirming it.
- To dismiss a system outright. A poor external result raises the question of fitness for your setting; it does not settle it.
Reach for something else instead
- Local validation — evaluate on your own population before clinical or operational use, which answers the question that matters to you.
- Prospective evaluation — measure the system in live use against outcomes, rather than retrospectively against recorded ones.
- -
Read more on the blog
- The pilot ran on data the production system will never seeEnterprise AI pilots are built on a curated slice, pre-cleaned, with limited users and manual review. Then production data arrives. That sequence explains the failure rate without invoking model capability at all.
- 72 seconds, or 30 minutes, and both are trialsTerritory 11 opens on the first subject this corpus has examined where the evidence is genuinely good. Registered trials, CONSORT-AI reporting, peer review, and effect sizes that still differ by a factor of twenty-five.
- The 12% was non-inferior, and P was 0.41MASAI is the best-evidenced AI deployment in medicine and the headline everyone quoted describes a result the trial did not claim.
- Nine subjects, and the model was almost never the blockerTerritory 10 closes. Across nine enterprise deployment subjects the binding constraint was organisational in every one, the evidence was commissioned in almost all of them, and the reconciling study exists nowhere.
Further reading
- Wong et al. (2021), External Validation of a Widely Implemented Proprietary Sepsis Prediction Model — the canonical case, and the source of the 7% figure.
- Raji et al. (2021), AI and the Everything in the Whole Wide World Benchmark — why a developer-selected evaluation cannot license a general claim.
Primary sources, listed so you can check the claims on this page rather than take them on trust.
Where people go wrong
- Accepting a vendor-run study on customer data as external. The evaluators matter as much as the population.
- Reporting discrimination without calibration, which hides whether the stated probabilities mean anything.
- Treating deployment scale as evidence. Hundreds of installations tell you about procurement, not performance.
At a glance
Where this sits
A starting point. Nothing needs to come before it.
Computed from the prerequisite graph, not assigned. How this works