Measurement Concentration
Research effort settling at the stage of a causal chain that is cheapest to instrument, which is rarely the stage that decides the outcome.
When not to use it
- Where the cheap stage is also the deciding stage, which does occur and makes the concentration harmless.
- As a criticism of individual studies, since the mechanism operates on the distribution rather than on any paper.
- Where funding explicitly targets the expensive link, which is the condition the concept describes the absence of.
Reach for something else instead
- Stage mapping — naming, alongside any claimed benefit, the stage at which it was measured.
- Targeted funding for decision-relevant links — the only intervention that changes the distribution rather than the individual studies.
- -
Read more on the blog
- 72 seconds, or 30 minutes, and both are trialsTerritory 11 opens on the first subject this corpus has examined where the evidence is genuinely good. Registered trials, CONSORT-AI reporting, peer review, and effect sizes that still differ by a factor of twenty-five.
- Phase I improved. Phase II did not.AI-designed drugs clear safety trials at well above industry rates. At the stage that tests whether a drug works, the sources contradict each other, and the more careful ones report no advantage at all.
- Adding the doctor to the model changed nothingA randomised trial found physicians did better with an LLM than with conventional resources. Its second comparison, reported in the same abstract, found the model alone did just as well as the model plus the physician.
- Clinicians override 49% to 96% of alertsFive years after the sepsis model was externally validated and found wanting, the sector-level evidence on clinical prediction alerts is process markers and no high-quality mortality signal.
Further reading
- Wong et al. (2021), External Validation of a Widely Implemented Proprietary Sepsis Prediction Model — discrimination measured widely, and the population that mattered evaluated only when somebody chose to.
- Raji et al. (2021), AI and the Everything in the Whole Wide World Benchmark — how the availability of a benchmark shapes what a field concludes it has measured.
Primary sources, listed so you can check the claims on this page rather than take them on trust.
Where people go wrong
- Reading a well-evidenced stage as evidence about the whole chain.
- Treating the absence of evidence at an expensive stage as evidence of no effect there.
- Attributing the gap to bias or negligence, when cost explains it without either.
At a glance
Where this sits
A starting point. Nothing needs to come before it.
Computed from the prerequisite graph, not assigned. How this works