Good evidence, and eight different failures
Territory 11 closes. This was the field chosen because its evidence is strong, and every subject produced a failure that stronger evidence would not have prevented.
TL;DR. This territory was entered because clinical AI has the best evidence infrastructure of any field this corpus has examined: registered trials, pre-specified protocols, blinding, ethics review, journals that publish nulls, and a regulator with statutory authority. Eight subjects later, every one produced a failure that more rigour would not have prevented. A twenty-five-fold effect range across settings. A P of 0.41 reported as a 12% reduction. A field rate disputed by 28 points. A benchmark inherited from medical education. A waitlist control. A framework governing modification rather than clearance. 1.3% of 232 studies recording the variable that mattered. And override rates of 49% to 96% defeating any model. One mechanism runs under all of them: measurement effort settles at the stage of a causal chain that is cheapest to observe, and that is never the stage that decides the outcome.
---
Status: synthesis. No new factual claims. Every figure appears in one of the eight Territory 11 articles or the two earlier clinical articles, with its own sourcing and caveats, linked where used. Evidence quality in this territory is the highest in the corpus, which is what makes the findings worth stating.
---
The eight
| Subject | The evidence | The failure |
|---|---|---|
| Scribes | Registered RCT, CONSORT-AI | Effect varies 25-fold by setting |
| Mammography | 105,915 women, three Lancet papers | P = 0.41 reported as a 12% reduction |
| Drug discovery | Registered trials throughout | Field rate disputed by 28 points |
| Diagnosis | Randomised, NCT06208423 | Benchmark inherited from education |
| Therapy | RCT in NEJM AI | Control group received nothing |
| Regulation | Statutory authority | Governs modification, not clearance |
| Dermatology | 232 pooled studies | 1.3% recorded skin type |
| Alerts | Systematic review, 44 implementations | 49 to 96% overridden |
Finding one: none of the eight is a shortage of rigour
Every study named above was competently conducted.
The mammography trial randomised a hundred thousand women, pre-specified three analyses and published its own late registration. The scribe trial used covariate-constrained randomisation and CONSORT-AI reporting. The therapy trial's limitations were identified in a letter published by its own journal.
Not one failure would have been prevented by a larger sample, tighter blinding, better statistics or stricter peer review.
They are failures of a different kind entirely: a setting effect, a transmission chain, an absent category register, an inherited task boundary, a comparator choice, a jurisdictional scope, an unrecorded variable, and a delivery channel.
Which is why this territory was worth entering. The corpus had spent five territories arguing that better evidence would settle contested questions. Clinical AI is where better evidence exists, and the argument can be tested rather than repeated.
The result is that better evidence produced better-specified uncertainty, which is genuine progress and is not what the argument promised.
Finding two: one mechanism runs under all of them
Measurement effort concentrates at the stage of a causal chain that is cheapest to observe.
In diagnosis, that stage is reasoning over an assembled case. Vignettes are abundant, scoreable and require no patients. Evidence assembly, which is the other half of diagnosis, requires a clinic and appears in almost no study.
In drug discovery, it is Phase I. Safety and pharmacokinetics are molecular properties with decades of training data, cleanly measurable, and reported consistently. Phase II, where biology decides, is reported at 40% and 68% by different analysts.
In agent and alert systems, it is discrimination. AUC can be computed retrospectively without deploying anything. Override rates require a deployment and a different research community, and the two literatures do not cite each other.
In scribes, it is minutes. Time-in-note is instrumented in the electronic record. Whether documentation is dreaded is not, and burnout moved thirteen points where the time saved might be sixteen minutes.
And in dermatology, it is accuracy. Aggregate accuracy needs no demographic column. Skin type needs a rater, and 1.3% of 232 studies used one.
The pattern is not laziness and it is not bias. Cheap-to-measure stages get measured. Expensive ones do not, regardless of which stage determines the result, and a field's evidence base ends up shaped like its instrumentation rather than like its causal structure.
Finding three: the chain stops before the outcome
Setting the stages out generalises the alert finding across the whole territory.
Stage one, the model performs: measured everywhere, in every subject, to high precision.
Stage two, the output reaches a person: rarely the bottleneck and rarely studied.
Stage three, the person acts on it: measured in one subject, where the answer was that 49% to 96% do not.
Stage four, the action changes: some evidence, mostly process markers such as time to antibiotics.
Stage five, the patient is better off: no high-quality evidence in any subject in this territory.
Not one of the eight measured a patient outcome. Diagnostic accuracy against a reference standard, documentation time, burnout scores, symptom scales, detection rates, interval cancers at two years, Phase II efficacy signals. All intermediate, all defensible, and none of them the thing.
And that is not a criticism of any individual study. Outcome trials take years, cost more, and require endpoints that intermediate measures were adopted precisely to avoid.
It is an observation about what the aggregate can support. A field with excellent evidence at stage one and none at stage five can say what its systems do and cannot say what they are worth.
Finding four: regulation cannot reach any of it
Medicine has the strongest regulator of any field this corpus has examined, and the failures sit outside its jurisdiction by construction.
A regulator assesses a submission: one product, one context of use, one sponsor. Every mechanism the FDA has built operates at that level, from predetermined change control plans to the credibility framework.
The eight failures are all above it. Synthesis across settings has no submission. Transmission after publication has no submission. Category-level aggregation, benchmark inheritance, comparator choice and unrecorded stratifiers have no submission either.
And two of them are below it. Override rates are a property of a hospital's alert configuration. EHR-embedded models positioned as clinical decision support frequently avoid clearance entirely, which is how a model deployed at hundreds of hospitals came to be externally validated only after the fact, by researchers who chose to run it.
The general form is that regulation is unit-scoped and the failures are relational, concerning how evidence compares, travels and aggregates. Those are properties of literatures, not of products.
Finding five: what good evidence did buy
A synthesis reporting only failures would misrepresent the territory, and three things were genuinely settled.
Vendor equivalence. The scribe trial compared two commercial products head to head and found remarkably similar performance and reception. No amount of vendor comparison could establish that, and it removes a procurement question organisations would otherwise spend months on.
Bounded claims. Ambient documentation saves time and reduces burnout, and both are now stated with sizes rather than adjectives. Mammography screening with AI is non-inferior on interval cancers at a 44.3% workload reduction. Those are defensible in a way they were not two years ago.
And known failure rates. Roughly one scribe note in fourteen contains fabricated content. That is a specific number a health system can design a review policy around, and it exists because somebody measured.
The pattern is that good evidence answers well-posed questions reliably. What it did not do is answer the questions the field is being cited for, because those questions sit at stages nobody instrumented.
What would settle each
| Subject | The missing measurement | Feasible now |
|---|---|---|
| Scribes | Review time offset, measured directly | Yes |
| Mammography | Mortality at ten years | No, requires time |
| Drug discovery | Published programme list per rate | Yes |
| Diagnosis | Augmentation comparison in real encounters | Yes |
| Therapy | Trial with an active control | Yes |
| Dermatology | Stratified re-evaluation of current products | Yes |
| Alerts | Alerts per patient-day against baseline | Yes, this quarter |
| All | One outcome trial, any subject | No, requires years |
Six of eight are feasible now, and five require no new methodology at all.
The two that require time are the two that matter most, which is the structural problem rather than a scheduling one.
And the alerts row is the cheapest measurement identified in this corpus. A rolling ratio of alerts to patient-days would have surfaced the pandemic finding in week one instead of in a retrospective study across 24 hospitals, and it is one query.
Testing the mechanism forward
The strongest objection to measurement concentration is that it explains everything after the fact. The defence is to state what it predicts before the evidence exists, so it can fail.
Four predictions, each falsifiable and none yet tested.
Where a task splits into a gathering half and a reasoning half, evaluations will measure the reasoning half. This should hold outside medicine: legal research, financial analysis, engineering diagnosis. If a major benchmark appears that requires a system to decide what information to obtain, and it becomes standard, the prediction fails.
Where a deployment has a human in the loop, the human's behaviour will be measured last and by a different community. Override, acceptance and edit rates should lag capability metrics by years in every domain. If an agent benchmark begins reporting acceptance rates alongside accuracy as standard, the prediction fails.
Where an outcome is deferred, intermediate endpoints will dominate and the deferred one will remain unmeasured for as long as intermediates are accepted. This predicts specifically that the first clinical AI outcome trial will come from a party with a regulatory reason to run one, not from a party curious about the answer.
And where a stratifying variable requires an extra step to record, it will be recorded in a small single-digit percentage of studies until a reporting guideline mandates it. This is checkable now in fields other than dermatology.
The general form is that the prediction is about cost rather than about subject. Any stage of any chain that requires new instrumentation will be under-measured relative to its importance, and the gap will be proportional to the instrumentation cost rather than to how much the stage matters.
Which is falsifiable by a single counterexample: a field that systematically measures its expensive stage while leaving a cheap one unmeasured. This corpus has not found one, and has not gone looking, which is worth stating as a weakness rather than as support.
What the territory does not show
That clinical AI does not work. Workload reductions of 44.3%, burnout falling thirteen points, sensitivity gains consistent across subgroups, and a Phase IIa efficacy signal in a disease where progression is rarely halted are all real.
That the evidence is bad. It is the best in this corpus, which is the premise of the whole territory.
That the researchers erred. Almost every failure here occurred downstream of a competent study or in a stage nobody was assigned to measure.
And that eight subjects characterise a field. They were selected for having evidence worth checking, which selects for well-studied areas.
What this does to the corpus's own argument
Five territories concluded that better disclosure and better measurement would settle contested questions. This one tested that, and the answer is partial.
Better measurement settled what it measured. Vendor equivalence, effect sizes, failure rates. Every one of those is now firmer than it was.
It did not settle the contested questions, because those live at stages the measurement did not reach, and the reason it did not reach them is that those stages are expensive.
Which qualifies the recommendation rather than refuting it. "Measure more" is right and incomplete. The operative version is measure the stage that decides, which is usually the one nobody has instrumented, and that costs more than measuring the stage that is already wired.
And it explains something the corpus had observed without accounting for. Across eleven territories, the recurring finding has been that the operative variable was structural rather than technical. This territory suggests why: technical variables sit at the instrumented stage and structural ones do not, so a field measuring what is cheap to measure will systematically produce evidence about capability and silence about everything else.
The counter-argument
Eight subjects chosen by one corpus is not a survey of clinical AI. They were selected for having checkable evidence, which selects for exactly the areas where the measurement critique applies, and a territory assembled that way will find measurement problems. This is the same selection objection the previous synthesis recorded and it has not been addressed.
Measurement concentration explains too much. Any field's evidence base can be described as concentrated at its cheapest stage after the fact, and a mechanism compatible with every observation is not doing explanatory work. The test is whether it predicts, and this article names five instances retrospectively.
Intermediate endpoints are not a failure of nerve. Medicine adopted surrogates because outcome trials take a decade, and demanding one for every AI intervention applies a standard the incumbent comparators never met. Diagnostic accuracy has been the accepted endpoint for imaging and laboratory testing for generations.
And the regulatory conclusion may be too pessimistic. A regulator that requires stratified reporting as part of a stated context of use would address the dermatology failure directly, and the FDA's credibility framework already turns on exactly that concept. The claim that regulation cannot reach these problems may describe the current instruments rather than the possible ones.
The short version
Eight subjects in the field with the best evidence infrastructure this corpus has examined, and eight different failures, none of which more rigour would have prevented: a twenty-five-fold effect range across settings, a P of 0.41 reported as a reduction, a field rate disputed by 28 points, a benchmark inherited from medical education, a waitlist control, a framework governing modification rather than clearance, 1.3% of 232 studies recording the variable that mattered, and 49% to 96% of alerts overridden.
One mechanism runs under all of them. Measurement effort settles at the stage of a causal chain that is cheapest to observe. Vignettes are abundant and clinics are not. Phase I is molecular and Phase II is biological. AUC is retrospective and override rates need a deployment. Minutes are instrumented and dread is not. Accuracy needs no demographic column and skin type needs a rater.
And the chain stops before the outcome. Stage one, the model performs: measured everywhere. Stage three, a person acts: measured once, at 49% to 96% not. Stage five, the patient is better off: no high-quality evidence in any subject in this territory.
Regulation cannot reach it. A regulator assesses one product for one use by one sponsor, and every failure here is relational: how evidence compares, travels and aggregates, which are properties of literatures rather than of products.
What good evidence did buy is real. Vendor equivalence no comparison could establish, effect sizes stated with numbers rather than adjectives, and a one-in-fourteen fabrication rate a health system can design around.
Which qualifies what this corpus has argued for five territories. Measure more is right and incomplete. Measure the stage that decides, and note that it is expensive precisely because nobody has wired it.
Common questions
What is the central finding of this territory? That eight subjects in the field with the strongest evidence infrastructure this corpus has examined each produced a failure that more rigour would not have prevented. The failures were a setting effect, a transmission chain, an absent category register, an inherited task boundary, a comparator choice, a jurisdictional scope, an unrecorded variable and a delivery channel. Not one would have been fixed by a larger sample, tighter blinding, better statistics or stricter peer review.
What connects them? Measurement effort concentrating at the stage of a causal chain that is cheapest to observe. Vignettes require no patients while evidence assembly requires a clinic. Phase I tests molecular properties with decades of training data while Phase II tests biology. Discrimination can be computed retrospectively while override rates require a deployment and a different research community. Time-in-note is instrumented in the record while whether documentation is dreaded is not. Aggregate accuracy needs no demographic column while skin type needs a rater.
Why does that matter more than any individual finding? Because it predicts where a field's evidence will be silent. A discipline measuring what is cheap to measure produces a body of evidence shaped like its instrumentation rather than like its causal structure, so it accumulates precision about capability and silence about consequence. That is not a criticism of any researcher; it is a property of how evidence bases assemble under cost constraints.
Did any study measure a patient outcome? No. Every figure in this territory is intermediate: diagnostic accuracy against a reference standard, documentation time, burnout scores, symptom scales, detection rates, interval cancers at two years, Phase II efficacy signals. All are defensible and none is the outcome that matters. The systematic review of implemented clinical prediction models found improved process markers, improved length of stay in two studies, one low-quality mortality result and no high-quality study showing a mortality difference.
Why can't regulation fix this? Because a regulator assesses a submission, which is one product, for one context of use, by one sponsor, and every mechanism the FDA has built operates at that level. The failures documented here are relational: synthesis across settings, transmission after publication, category-level aggregation, benchmark inheritance, comparator choice and unrecorded stratifiers. Those are properties of literatures rather than of products, and two of them sit below the regulator instead, since EHR-embedded models positioned as clinical decision support frequently avoid clearance.
What did good evidence actually settle? Three things worth naming. Vendor equivalence, where a randomised trial compared two commercial scribes head to head and found remarkably similar performance, which no vendor comparison could establish. Bounded claims, so that time savings, burnout reduction and screening non-inferiority at a 44.3% workload reduction are stated with sizes rather than adjectives. And known failure rates, including roughly one scribe note in fourteen containing fabricated content, which is a number a health system can design a review policy around.
What would settle the open questions? Six of eight are feasible now and five require no new methodology: measuring review time offset directly, publishing the programme list behind each drug success rate, running the augmentation comparison in real encounters, running a therapy trial with an active control, re-evaluating current dermatology products on a stratified benchmark, and computing alerts per patient-day against a rolling baseline. The two that require time are mortality outcomes in screening and one outcome trial in any subject, which are the two that matter most.
What is the strongest objection to this synthesis? That measurement concentration explains too much. Any field's evidence base can be described after the fact as concentrated at its cheapest stage, and a mechanism compatible with every observation is not doing explanatory work. The test is whether it predicts rather than describes, and this article names its instances retrospectively. A second objection is the selection one: eight subjects chosen for having checkable evidence selects for the areas where a measurement critique applies.
Sources
Primary documents only. Where a claim rests on a single report, the entry says so.
- The eight Territory 11 articles Artifipedia This is a synthesis and makes no new factual claims. Every figure appears in one of the eight articles, or in the two earlier clinical articles on medical devices and the sepsis model, each with its own sourcing and caveats. Evidence quality in this territory is the highest in the corpus, which is what makes the failures worth recording.
Related articles
- Phase I improved. Phase II did not.AI-designed drugs clear safety trials at well above industry rates. At the stage that tests whether a drug works, the sources contradict each other, and the more careful ones report no advantage at all.
- One trial, a waitlist control, and a letterThe best evidence for AI mental health support is a single randomised trial of a purpose-built clinical tool. Its own journal published three methodological objections, and almost nobody uses the thing that was tested.
- 232 studies, and 1.3% recorded skin typeA systematic review found AI detecting skin cancer at 90% accuracy across 232 studies. Almost none of those studies recorded who the patients were, and the ones that checked found performance dropping to chance.
- Plus 48% with the tool, minus 17% without itTerritory 12 opens on education, where the randomised evidence is unusually good and points in opposite directions depending on whether the test allows the tool.