Adding the doctor to the model changed nothing
A randomised trial found physicians did better with an LLM than with conventional resources. Its second comparison, reported in the same abstract, found the model alone did just as well as the model plus the physician.
TL;DR. A randomised controlled trial published in Nature Medicine, registered as NCT06208423, found physicians using an LLM scored 6.5 percentage points higher than those using conventional resources, 95% CI 2.7 to 10.2, P < 0.001. The same abstract reports a second comparison: between LLM-augmented physicians and the LLM alone, the difference was −0.9 points, 95% CI −9.0 to 7.2, P = 0.8. Adding a physician to the model produced no measurable change. A separate trial found physicians with LLM assistance did not significantly outperform those without, while the model alone outscored both groups. And a systematic review of 30 studies covering 19 models and roughly 4,762 cases found LLM diagnostic accuracy ranging from 25% to 97.8% and generally below physician accuracy. The vignette results and the pooled results point opposite ways.
---
Status: strong individual trials, contested aggregate. The primary sources are registered randomised trials in Nature Medicine and a six-experiment study in Science, alongside a systematic review pooling 30 studies. The trials and the review disagree, and the article's subject is why. Nothing here is clinical advice.
---
The comparison that gets quoted
*A randomised controlled trial, registered as NCT06208423 and published in Nature Medicine, gave physicians either an LLM or conventional resources such as reference databases and search, on complex clinical vignettes.*
Physicians using the LLM scored significantly higher: a mean difference of 6.5 percentage points, 95% CI 2.7 to 10.2, P below 0.001.
They also spent longer, by a mean of 119.3 seconds per case, 95% CI 17.4 to 221.2.
That is a clean result and it is the one in circulation: LLM assistance improves physician management reasoning against conventional tools.
The comparison that does not
The same abstract reports a second contrast.
Between LLM-augmented physicians and the LLM operating alone, the difference was −0.9 percentage points, 95% CI −9.0 to 7.2, P = 0.8.
No measurable difference. The model working by itself performed as well as the model working with a doctor.
And a separate randomised trial found something sharper. Internal medicine physicians given ChatGPT as a diagnostic aid on case vignettes did not significantly outperform physicians using conventional resources. The chatbot's suggestions sometimes helped and sometimes distracted, producing no net improvement in diagnostic reasoning scores. In the same study, the LLM answering alone outscored both physician groups.
Two trials, and in both the human contribution to the combined performance was not detectable.
That is a substantially more consequential finding than the headline one, because it bears directly on how such systems would be deployed, and it is the number nobody repeats.
Why the trials and the review disagree
A systematic review pooling 30 studies across 19 different LLMs and roughly 4,762 cases found primary diagnosis accuracy ranging from 25% to 97.8%, and generally below physician accuracy in most scenarios.
That directly contradicts the vignette headlines, and the resolution is in what each measured.
The strongest LLM results come from curated material. Clinicopathological conference cases are teaching artefacts: the diagnosis is known, the relevant findings are present, the narrative has been assembled by someone who knew the answer, and the case was selected for being instructive.
On those, one reasoning model reached exact or very close diagnostic accuracy in 88.6% of cases against GPT-4's 72.9%.
The pooled review covers a wider and messier range of tasks, including specialties where the input is images, cases where the presentation is ambiguous, and settings where the model receives what a clinician would actually have.
So the vignette is not the job. A clinicopathological conference case supplies the information; a patient does not. The work of medicine includes deciding what to ask, what to examine, what to order, and what to disregard, and none of that is tested by a case whose relevant facts were selected in advance.
This is construct validity in its most consequential clinical form. The vignette measures reasoning over assembled evidence. Diagnosis is reasoning plus evidence assembly, and the second half is invisible to the benchmark.
The result that partly survives the objection
Worth stating carefully, because one study went beyond vignettes.
*A six-experiment study published in Science tested a reasoning model against physician and prior-model baselines, and its sixth experiment used 76 actual emergency department cases* rather than teaching material.
At initial triage, the model reached exact or very close diagnostic accuracy in 67.1% of cases, against two expert attending physicians at 55.3% and 50.0%.
That is real clinical material and the model was ahead.
Three qualifications belong with it. Seventy-six cases is small, and two attending physicians is a very narrow human comparison. The model received the case as text, which means somebody had already recorded the history and findings. And "initial triage" is one of three touchpoints tested, which makes it a stage-specific result of exactly the kind the drug discovery article found elsewhere in this territory.
What it does establish is that the advantage is not purely an artefact of curated cases. The gap narrows on real material and does not vanish.
What the augmentation finding means
The two trials agreeing that the physician added nothing measurable is the result with deployment consequences, and it admits several readings.
The optimistic reading: the model is good enough that supervision is redundant on this task, and the physician's time could be spent elsewhere.
The cautious reading: the trial measured score on vignettes, and a physician's contribution to a real encounter includes examination, history-taking, judgement about what the patient did not say, and responsibility for the decision. None of that appears in a vignette score, so finding no contribution measures the instrument rather than the clinician.
The uncomfortable reading: physicians reviewing model output may exhibit the same pattern documented elsewhere in this corpus, where a plausible suggestion anchors judgement. The finding that suggestions "sometimes helped but also introduced distractions" is consistent with that.
This corpus favours the second reading and cannot demonstrate it. What the trials measured is a vignette score, and a vignette score is where a physician's distinctive contribution is least visible.
Which makes the honest summary uncomfortable in both directions. The trials do not show that doctors are unnecessary. They show that on the narrow task the trials measured, the doctor's presence did not move the number, and that task was chosen because it is measurable rather than because it is the job.
What is missing from all of it
No trial in this literature measures a patient outcome.
Every figure here is diagnostic accuracy against a reference standard: a discharge diagnosis, a conference case answer, or a scored management plan. None measures whether the patient did better.
That distinction is not pedantic in medicine. A more accurate diagnosis that arrives at the same treatment changes nothing. A diagnosis that is technically correct and triggers a cascade of confirmatory testing can cause harm. The endpoint that matters is downstream of everything measured here.
And the corpus has seen this exact gap before. The mammography trial measured detection, workload and interval cancers, with mortality unmeasured and unmeasurable at two years. The sepsis model was evaluated on discrimination and failed on the population that mattered.
Diagnostic accuracy is a proxy for benefit, it is far easier to measure than benefit, and the entire literature is built on it.
What each study measured, side by side
The disagreement dissolves once the tasks are set against each other.
| Study | Material | Comparison | Result |
|---|---|---|---|
| Nature Medicine RCT | Complex vignettes | Physician + LLM vs conventional | +6.5 points |
| Same trial, secondary | Complex vignettes | Physician + LLM vs LLM alone | −0.9, P = 0.8 |
| Second RCT | Case vignettes | Physician + LLM vs physician | No significant gain |
| Science, exp. 1 to 5 | Conference cases | Model vs model, model vs physician | 88.6% vs 72.9% |
| Science, exp. 6 | 76 real ED cases | Model vs two attendings | 67.1% vs 55.3%, 50.0% |
| Systematic review | 30 studies, mixed | Pooled LLM vs clinicians | 25 to 97.8%, generally below |
Read the material column and the results order themselves.
Curated teaching cases produce the largest model advantage. Real clinical material narrows it. Pooling across many task types reverses it.
That ordering is exactly what a construct-validity problem produces, and it is more informative than any single row. The model's advantage is inversely proportional to how much of the diagnostic work was done before the case reached it.
And the two rows worth reading together are the second and the sixth. The physician added nothing measurable on vignettes, and the pooled evidence across mixed real tasks puts models generally below physicians. Both can hold if the vignette is the task where the model is strongest and the clinician is least visible.
What would settle it
Three studies, in ascending order of difficulty, and only one is expensive.
Report the augmentation comparison on real encounters. The finding that physicians added nothing was measured on vignettes. The same comparison in a clinic, with the model receiving what the clinician receives, is a straightforward trial design and nobody has run it.
Measure the assembly half. Every study here hands the model a written case. A design where the system must decide what to ask and what to order, scored against a clinician doing the same, would test the part currently invisible. Simulated patient interviews have been used this way and are the nearest existing approach.
And run one trial with a patient outcome. Not diagnostic accuracy, but whether the treatment changed and whether the patient did better. This is the expensive one, requires years, and is the only design that answers the question the literature is being cited for.
The first two require no new methodology. They require somebody to run the comparison on the harder task rather than the measurable one, which is the pattern this corpus has found in every territory and which appears here in a field with registered trials and peer review.
Three things this establishes
The second comparison is the important one and it does not travel. A 6.5-point improvement over conventional resources circulates. A P of 0.8 between the model alone and the model with a physician sits in the same abstract, and it is the number that bears on how these systems would actually be deployed.
Curated cases and pooled studies point opposite ways, and both are right. A reasoning model reaching 88.6% on conference cases and a review finding accuracy generally below physicians describe different tasks. The vignette supplies the information; the encounter is where it is gathered.
And nothing measures whether patients are better off. Every result in this literature is accuracy against a reference standard, which is a proxy chosen because it is measurable, in a field where the difference between accuracy and benefit is the whole of clinical medicine.
What it does not establish
That LLMs are not useful diagnostically. A 67.1% against 55.3% and 50.0% result on real emergency cases is a genuine finding, and the second-opinion use case is the one the Science authors emphasise.
That physicians add nothing. Two trials found no measurable contribution on a vignette score, which is the task least suited to detecting what a clinician does.
That the systematic review is decisive. Pooling 30 studies across 19 models and a range of tasks produces a very wide band, and heterogeneity that large limits what a pooled figure can support.
And nothing about any deployment. No system described here has been evaluated in routine clinical use with patient outcomes.
What is unresolved
Whether the augmentation null holds outside vignettes. The finding that matters most rests on the task where it is least testable, and a trial in real encounters would settle it.
What the physician contributes that a vignette cannot capture. Named here as examination, history-taking and responsibility, and unmeasured.
Whether accuracy translates to outcomes. The question the literature does not ask, requiring trials with clinical endpoints and years of follow-up.
And whether reasoning models change the picture again. The 88.6% against 72.9% gap between model generations was large, and the comparison baselines in most of this literature are already outdated.
The pattern this corpus keeps meeting
Territory 7 established that physical automation succeeded wherever the specification was negotiable, and named five forms, one of which was environment engineering: the world was rearranged so the task became tractable.
Warehouse robots did not learn to walk. The floor was flattened, the shelves were standardised, and the walking was deleted.
The same move appears here at a different layer.
A clinicopathological conference case is an engineered environment for a diagnostic system. The history has been taken, the examination recorded, the imaging ordered and reported, the irrelevant findings pruned, and the narrative arranged by someone who knew the answer. What remains is the reasoning step.
And the reasoning step is what the model does well, which is why the results are strong and why they narrow on real emergency cases where less of that preparation has happened.
The distinction worth holding is that in robotics the engineering was visible and deliberate. Somebody flattened the floor, and the corpus could count what had been done.
In evaluation the engineering is upstream and invisible, because the case arrives already assembled and the assembly was performed by the medical education system decades before anyone thought to benchmark a model on it. Nobody prepared the environment for the AI. The environment was already prepared, for teaching.
Which means the benchmark is not measuring what its users think and is not anybody's fault. Conference cases were built to test whether a trainee can reason from evidence, on the reasonable assumption that gathering evidence was separately assessed. A model taking the same test inherits the assumption without inheriting the separate assessment.
The generalisation is worth stating. Where a benchmark is inherited from human education, it will test the part of a task that education isolated for testing, and that isolation was designed around what humans find hard rather than around what the whole job requires.
And humans find reasoning hard and gathering routine. For a model the difficulty runs the other way, so a benchmark built for the human difficulty curve systematically flatters a system with the opposite one.
The counter-argument
Treating the LLM-alone comparison as the headline overreaches. It was a secondary contrast in a trial powered for the primary one, its confidence interval spans −9.0 to 7.2, and a wide interval containing zero is weak evidence of equivalence rather than evidence of no difference. This article makes a null result carry more than it can.
The vignette objection is applied selectively. Physicians in these trials also worked from vignettes, so both sides faced the same artificial task and the comparison is internally fair. Criticising the benchmark for not being clinical practice does not favour either party, and the article uses it only in one direction.
The systematic review's range is too wide to conclude from. Accuracy from 25% to 97.8% across 19 models and 4,762 cases describes heterogeneity rather than a finding, and reading "generally below physician accuracy" from that band is a summary the underlying variance may not support.
And the outcomes objection would disqualify most of medicine. Diagnostic accuracy has been the accepted intermediate endpoint for imaging, laboratory testing and clinical examination for decades. Requiring outcome trials for LLM diagnosis sets a bar the incumbent comparator never cleared, which is the same asymmetry this corpus criticises when applied to other technologies.
The short version
*A randomised trial in Nature Medicine, registered as NCT06208423, found physicians using an LLM scored 6.5 points higher than those using conventional resources, 95% CI 2.7 to 10.2, P below 0.001, while spending 119.3 more seconds per case.*
The same abstract reports the second comparison: LLM-augmented physicians against the LLM alone, a difference of −0.9 points, 95% CI −9.0 to 7.2, P = 0.8. Adding a physician produced no measurable change. A separate trial found physicians with an LLM did not significantly outperform those without, while the model alone outscored both groups.
Against that, a systematic review of 30 studies across 19 models and roughly 4,762 cases found accuracy from 25% to 97.8%, generally below physician accuracy.
Both are right because they measure different things. Curated conference cases, where a reasoning model reached 88.6% against GPT-4's 72.9%, supply the information, the answer is known, and the case was chosen for being instructive. A patient supplies none of that, and deciding what to ask, examine and order is the half a vignette cannot test.
One result partly survives the objection. On 76 actual emergency department cases, a reasoning model reached 67.1% at initial triage against two attending physicians at 55.3% and 50.0%. Small, narrow, text-only, and real.
And nothing in this literature measures a patient outcome. Every figure is accuracy against a reference standard, which is a proxy chosen for being measurable in the field where the gap between accuracy and benefit is the entire discipline.
Common questions
Do LLMs outperform doctors at diagnosis? It depends entirely on the task. On curated clinicopathological conference cases, a reasoning model reached exact or very close diagnostic accuracy in 88.6% of cases against GPT-4's 72.9%, and on 76 real emergency department cases the same model reached 67.1% at initial triage against two attending physicians at 55.3% and 50.0%. Against that, a systematic review pooling 30 studies across 19 models and roughly 4,762 cases found accuracy ranging from 25% to 97.8% and generally below physician accuracy in most scenarios.
What did the randomised trial actually find? Two things. Physicians using an LLM scored 6.5 percentage points higher than those using conventional resources, with a 95% confidence interval of 2.7 to 10.2 and P below 0.001, while spending 119.3 seconds longer per case. And in a second comparison reported in the same abstract, the difference between LLM-augmented physicians and the LLM operating alone was −0.9 points, 95% CI −9.0 to 7.2, P = 0.8. The model by itself performed as well as the model with a doctor.
Why does that second finding matter? Because it bears on deployment in a way the first does not. If augmentation improves on conventional tools but adds nothing to the model working alone, the question of what role the clinician plays becomes practical rather than theoretical. It is also the number that does not circulate, despite appearing in the same abstract as the one that does.
Should that be read as doctors being unnecessary? No, and the article gives three readings. The trial measured a vignette score, and a physician's contribution to a real encounter includes examination, history-taking, judgement about what a patient did not say, and responsibility for the decision, none of which a vignette captures. Finding no measurable contribution on that task may measure the instrument rather than the clinician. A less comfortable possibility is that reviewers anchor on plausible model output, which is consistent with the finding that suggestions sometimes helped and sometimes distracted.
Why do the trials and the systematic review disagree? Because the strongest results come from curated material. A clinicopathological conference case is a teaching artefact: the diagnosis is known, the relevant findings are present, the narrative was assembled by someone who knew the answer, and the case was chosen for being instructive. The pooled review covers a wider and messier range including image-based specialties and ambiguous presentations. The vignette measures reasoning over assembled evidence; diagnosis is reasoning plus evidence assembly.
Is there any result on real clinical material? Yes, and it is the strongest evidence in the subject. A six-experiment study in Science included 76 actual emergency department cases, where a reasoning model reached 67.1% exact or very close diagnostic accuracy at initial triage against two attending physicians at 55.3% and 50.0%. The qualifications are that 76 cases is small, two physicians is a narrow comparison, the model received the case as text after somebody had recorded the history, and initial triage was one of three touchpoints tested.
What does none of this measure? Patient outcomes. Every figure in this literature is diagnostic accuracy against a reference standard: a discharge diagnosis, a conference case answer, or a scored management plan. None measures whether patients did better. A more accurate diagnosis leading to the same treatment changes nothing, and a technically correct diagnosis triggering confirmatory testing can cause harm.
What is the strongest objection to this article? That it makes a null result carry too much. The LLM-alone comparison was secondary in a trial powered for the primary one, and a confidence interval spanning −9.0 to 7.2 is weak evidence of equivalence rather than evidence of no difference. A second objection is that the outcomes complaint would disqualify most of medicine, since diagnostic accuracy has been the accepted intermediate endpoint for imaging, laboratory testing and clinical examination for decades.
Sources
Primary documents only. Where a claim rests on a single report, the entry says so.
- GPT-4 assistance for improvement of physician performance on patient care tasks: a randomized controlled trial Nature Medicine, registered NCT06208423 Both comparisons in the same abstract: physicians with the LLM scoring 6.5 points higher than with conventional resources, 95% CI 2.7 to 10.2, P below 0.001; and LLM-augmented physicians against the LLM alone at -0.9 points, 95% CI -9.0 to 7.2, P = 0.8.
- Could AI Surpass Doctors at Clinical Reasoning? Inside Precision Medicine, on the Science study The six-experiment study: 88.6% against GPT-4's 72.9% on clinicopathological conference cases, and 67.1% at initial triage on 76 actual emergency department cases against two attending physicians at 55.3% and 50.0%.
- Comparing Diagnostic Accuracy: LLMs vs. Physicians IntuitionLabs The systematic review pooling 30 studies across 19 models and roughly 4,762 cases, finding accuracy from 25% to 97.8% and generally below physician accuracy, and the second randomised trial where LLM assistance produced no net improvement while the model alone outscored both physician groups.
Further reading
The primary literature behind the claims above, drawn from the concept entries this post links to, so a claim carries the same source here as it does there.
- Raji et al. (2021), AI and the Everything in the Whole Wide World Benchmark — why general benchmarks cannot carry general claims. :: https://arxiv.org/abs/2111.15366 Construct Validity
- Bowman & Dahl (2021), What Will it Take to Fix Benchmarking in Natural Language Understanding? — what a benchmark must satisfy to support inference. :: https://arxiv.org/abs/2104.02145 Construct Validity
- Parasuraman & Riley (1997), Humans and Automation: Use, Misuse, Disuse, Abuse — the paper that named the failure modes. :: https://journals.sagepub.com/doi/10.1518/001872097778543886 Automation Bias
- Geirhos et al. (2020), Shortcut Learning in Deep Neural Networks — why systems fail on the variation a reviewer is least primed to catch. :: https://arxiv.org/abs/2004.07780 Automation Bias
Related articles
- The 12% was non-inferior, and P was 0.41MASAI is the best-evidenced AI deployment in medicine and the headline everyone quoted describes a result the trial did not claim.
- Phase I improved. Phase II did not.AI-designed drugs clear safety trials at well above industry rates. At the stage that tests whether a drug works, the sources contradict each other, and the more careful ones report no advantage at all.
- Plus 48% with the tool, minus 17% without itTerritory 12 opens on education, where the randomised evidence is unusually good and points in opposite directions depending on whether the test allows the tool.
- 99.98% is a tracking rate, and the field is animationNVIDIA's motion controller is a real advance that routes around the data problem Territory 7 identified. The number attached to it measures something narrow, and it gets weaker the further it travels from the field it was measured in.