Aggregate Evidence Gap
Where every individual study is rigorously produced and the field-level number is unreliable, because quality control attaches to the artefact and nobody owns the sum.
When not to use it
- Where a maintained register with stated inclusion criteria exists, which is the condition the concept describes the absence of.
- Where the aggregate is itself the subject of a systematic review with published methods.
- As a reason to distrust individual studies, which are usually the reliable part.
Reach for something else instead
- Published category register — the list, with inclusion criteria, which makes every downstream rate checkable.
- Denominator linking — requiring any published rate to name the population it was computed over.
- -
Read more on the blog
- 61% once, 25% eight times runningEvery agent benchmark score you have seen is a single-attempt number. The metric that measures whether an agent does the same thing twice tells a different story, and almost nobody reports it.
- Phase I improved. Phase II did not.AI-designed drugs clear safety trials at well above industry rates. At the stage that tests whether a drug works, the sources contradict each other, and the more careful ones report no advantage at all.
- The framework that exists produced 1.6%The FDA has two AI tracks. One is final, has authorised over 1,350 devices, and is the regime under which almost none of them cite a trial. The other missed its own deadline five weeks ago.
- One trial, a waitlist control, and a letterThe best evidence for AI mental health support is a single randomised trial of a purpose-built clinical tool. Its own journal published three methodological objections, and almost nobody uses the thing that was tested.
Further reading
- Wong et al. (2021), External Validation of a Widely Implemented Proprietary Sepsis Prediction Model — a widely deployed system whose aggregate performance required an independent study nobody was obliged to run.
- Raji et al. (2021), AI and the Everything in the Whole Wide World Benchmark — how a category's composition determines what its aggregate can support.
Primary sources, listed so you can check the claims on this page rather than take them on trust.
Where people go wrong
- Treating a field-level rate as inheriting the reliability of the peer-reviewed studies beneath it.
- Comparing two aggregate figures without checking whether either published a denominator.
- Assuming a well-regulated field has good field-level data, when regulation is almost always unit-scoped.
At a glance
Where this sits
A starting point. Nothing needs to come before it.
Computed from the prerequisite graph, not assigned. How this works