Home/Blog/Evaluation & evidence/72 seconds, or 30 minutes, and both are trials

72 seconds, or 30 minutes, and both are trials

Territory 11 opens on the first subject this corpus has examined where the evidence is genuinely good. Registered trials, CONSORT-AI reporting, peer review, and effect sizes that still differ by a factor of twenty-five.

TL;DR. Ambient AI scribes are the first AI deployment this corpus has examined where the evidence base is registered trials rather than vendor surveys. A three-arm randomised trial at UCLA, registered as NCT06792890, written to SPIRIT-AI and reported under CONSORT-AI, compared two commercial scribes against usual care and found roughly 7% improvement in burnout scores with remarkably similar performance across both vendors. A study across six health systems found burnout falling from 51.9% to 38.8% in 30 days. And the time savings range enormously: 72.6 seconds per encounter in a Stanford emergency department cohort of 10,336 encounters, against about 30 minutes per provider per day at UW Health and 16 minutes per eight hours across five academic centres. The spread is real variation, not measurement failure, which is what good evidence looks like. And roughly 7% of notes contain fabricated content.

---

Status: unusually well evidenced, and the evidence disagrees. Sources include a registered randomised trial published in NEJM AI, a quality improvement study in JAMA Network Open, a peer-reviewed retrospective cohort, and a multicentre study of around 1,800 clinicians. Evidence grade is stated per finding below, because the differences between a randomised trial, a cohort and a pre-post survey are the article's subject as much as the numbers are.

---

Why this subject is different

Every previous territory in this corpus has had a measurement problem. Enterprise deployment evidence was commissioned by parties selling the remedy. Information-ecosystem figures came from detection vendors. Physical economics rested on disclosure that mostly did not exist.

This one has registered trials.

The UCLA study was a parallel three-arm pragmatic randomised trial, with physicians assigned one-to-one-to-one by covariate-constrained randomisation balancing on time-in-note, baseline burnout score and clinic days per week, to one of two commercial scribes or a usual-care control, running from 4 November 2024 to 3 January 2025.

It was developed in accordance with SPIRIT-AI, registered on ClinicalTrials.gov as NCT06792890, and reported under CONSORT-AI. It states its own protocol limitation openly: institutional urgency to test the tools during a brief contractual window precluded timely pre-registration, with the finalised protocol submitted on 9 December 2024 and published on 27 January 2025.

That last paragraph is what a well-run study looks like. The registration was late, the authors say so, and a reader can weigh it.

Which makes this territory a test of something this corpus has repeatedly implied: that better evidence produces clearer answers. It partly does, and the clarity arrives in an unexpected place.

The time findings, and their spread

Four measurements of documentation time, four very different numbers.

72.6 seconds reduction per encounter, in a Stanford emergency department retrospective cohort of 10,336 encounters, using mixed-effects linear models with physician random intercepts, 95% confidence interval 63.8 to 81.4 seconds, P below .001.

About 30 minutes less documentation per provider per day, in a randomised trial at UW Health, after which the system was rolled out to roughly 800 clinicians across Wisconsin and Illinois.

About 16 minutes saved per eight hours of patient care, plus 13 fewer minutes in the record overall, across roughly 1,800 clinicians at five academic medical centres, published in early 2026.

And measurable time savings in the UCLA randomised trial, with both vendors performing similarly.

The range from 72.6 seconds per encounter to 30 minutes per day is not a contradiction. An emergency physician seeing twenty patients a shift at 72.6 seconds each saves about 24 minutes. The units differ, the settings differ, and the reconciliation is arithmetic rather than dispute.

What does not reconcile as easily is 16 minutes per eight hours against 30 minutes per day, and the larger study reports the smaller effect, which is the ordinary direction as sample size rises.

The finding that is more consistent than time

Burnout moved more reliably than minutes did, and that is the interesting result.

A quality improvement study across six academic and community health systems, using pre-intervention and 30-day post-intervention surveys of 131 ambulatory clinicians, found the proportion experiencing burnout falling from 51.9% to 38.8%, with an odds ratio of 0.26.

The UCLA randomised trial found roughly 7% improvement in burnout scores, which its own authors describe as a secondary endpoint needing confirmation in larger multicentre trials.

The UW Health trial reported a clinically meaningful reduction.

So a thirteen-point drop in the proportion burned out sits alongside time savings that may be sixteen minutes. Those are not proportionate, and the mismatch is the finding rather than a problem with it.

The plausible reading is that the benefit is not time. Documentation is not merely long; it is the specific aversive task that follows the working day. Removing a task people dread is not the same as removing minutes, and a measure of minutes will understate it.

Which is a scope question rather than a discrepancy. Time-in-note measures duration. Burnout instruments measure something the duration was a proxy for, and the proxy is weaker than assumed.

What the technology actually does wrong

Roughly 7% of notes contain fabricated content, which is approximately one note in fourteen.

Physical examination documentation is reported as unreliable, which is the section a scribe cannot observe: the model hears the consultation and does not see the examination.

Physician review is mandatory rather than optional, and the time required for that review partially offsets the time saved by ambient capture.

And no vendor accepts clinical liability.

That last point is the one with the sharpest structure. The system drafts, the physician signs, and the signature carries the legal weight. The efficiency accrues to the workflow and the liability stays entirely with the clinician, which is cost externality in a setting where the cost is professional and personal rather than financial.

Related research is examining exactly where this bites. One study compares clinicians' modification of hedging language between AI drafts and final notes, and another analyses clinician edits to ambient drafts. Both are asking what the reviewer had to fix, which is the measurement the time studies do not capture.

The finding about training, which nobody quotes

The five-centre study that found the smallest time savings also found something more actionable.

Use was inconsistent, and the biggest benefits went to clinicians who were coached on how to work with the tool rather than left to work it out alone.

That is a deployment finding rather than a technology finding, and it is the same conclusion Territory 10 reached across nine subjects: the operative variable was organisational.

It also explains part of the spread. A trial where participants were recruited by departmental nomination and leadership referral, in a system motivated to test the tool, is measuring a supported deployment. A multicentre study of 1,800 clinicians is measuring something closer to ordinary conditions, and it found less.

Which is the standard relationship between effect size and sample breadth, and is a reason to weight the smaller number, not the larger one.

What good evidence bought

It is worth saying plainly what changed by having trials.

Vendor equivalence. The randomised trial found remarkably similar performance and reception across two distinct commercial platforms. No vendor comparison could have established that, and the finding removes a question buyers would otherwise spend months on.

A confirmed direction with a bounded size. Time savings are real and modest. Burnout improvement is real and larger than the time savings explain. Both statements are now defensible in a way they were not in 2024.

And a known failure mode with a number attached. A 7% fabrication rate is a specific claim that can be designed around: it makes review mandatory rather than advisory, and it tells a health system what its review policy is for.

What good evidence did not buy is a single effect size, and the reason is that there is not one. 72.6 seconds in an emergency department and 30 minutes in ambulatory care are both correct, because the settings genuinely differ.

This corpus has spent several territories arguing that conflicting figures usually indicate a definitional problem. Here they do not. Sometimes a spread is the world, and distinguishing that case from the other one is the harder skill.

The evidence, graded

Stating the design alongside each finding is the whole point of a subject with real trials, and it is rarely done in coverage.

FindingDesignSampleGrade
~7% burnout improvement, vendors equivalentRandomised, 3-arm238 analysedLevel I
~30 min/day documentation savedRandomisedUW HealthLevel I
16 min per 8 hours, 13 min in recordMulticentre observational~1,800Level II
51.9% to 38.8% burnout in 30 daysPre-post survey, no control131Level II
72.6 sec per encounterRetrospective cohort10,336 encountersLevel III

Read the grade column against the effect column and a pattern appears.

The two randomised findings are the modest ones on time and the cautious one on burnout. The dramatic burnout figure is the design with no control group. The most precise time figure, with the tightest confidence interval and the largest encounter count, is the retrospective cohort, which is the weakest design and the strongest statistics.

None of that means the weaker designs are wrong. It means the confidence a reader should attach differs by row, and the rows are quoted interchangeably everywhere.

And it produces a useful ordering. If only one claim survives: ambient scribes reduce documentation time modestly and reduce burnout, both vendors work about equally, and one note in fourteen contains something that was not said. That sentence is supported at Level I or by direct measurement throughout, and it is considerably duller than the coverage.

What this does to the corpus's own thesis

Territory 10 closed on the finding that its evidence was almost entirely commissioned, with four reliable anchors across nine subjects. This subject is the counterfactual.

So the question worth asking directly: did better evidence change the conclusion?

Partly, and not in the expected direction.

It confirmed a benefit that vendor material also claimed. Time savings and burnout reduction were in the marketing before they were in the trials, and the trials found them. This corpus has repeatedly found vendor claims collapsing under measurement, and here one did not.

It removed a question rather than answering one. Vendor equivalence is the sort of finding only a randomised comparison produces, and its practical effect is to make a procurement decision smaller.

It bounded the size. The claims survive and are more modest than the enthusiasm, which is the ordinary result of measurement and is worth stating as a success rather than a debunk.

And it produced a disagreement between two real measures, time and burnout, that no amount of further precision on either will resolve, because they are measuring different things and only one of them is what anyone cares about.

Which qualifies something this corpus has been implying for five territories. The recurring finding has been that better disclosure and better measurement would settle contested questions. Here they were available and the contested question moved rather than closing, from how much time is saved to what the time was standing in for.

That is progress and it is not resolution, and a corpus arguing for measurement should be honest that this is what measurement usually delivers.

Three things this establishes

Registered trials changed what can be claimed, and did not collapse the range. SPIRIT-AI protocols, CONSORT-AI reporting and randomisation produced defensible statements about direction, vendor equivalence and failure rate. They did not produce one number, because the effect genuinely varies by setting.

The benefit and the measurement are different things. Burnout fell thirteen points in thirty days while time savings may be sixteen minutes. A measure of minutes was standing in for the experience of a dreaded task, and it turns out to be a poor proxy.

And the liability did not move. The scribe drafts, roughly one note in fourteen contains fabricated content, review is mandatory, review time offsets capture savings, and no vendor accepts clinical responsibility. The efficiency is distributed and the exposure is not.

What it does not establish

That ambient scribes do not work. The direction is consistent across randomised, cohort and survey evidence, which is a stronger position than most AI deployments have.

That the burnout effect is durable. The strongest burnout figure comes from a 30-day pre-post survey, which is a design that captures novelty alongside benefit.

That the fabrication rate is stable. A single reported figure across an evolving product category is a snapshot, and no series exists.

And nothing about patient outcomes. The trials measured documentation time and clinician burnout. Whether note quality affects care is a different question that these studies did not ask.

What is unresolved

Whether the burnout improvement persists past thirty days. No published follow-up extends far enough, and the pre-post design is the weakest in the evidence set on exactly the finding that is strongest.

What review actually costs. Review time is stated to partially offset capture savings and nobody has measured the offset directly.

What the fabrications are. A 7% rate is a frequency without a taxonomy, and whether the errors are clinically trivial or material is the question that determines the risk.

And whether coaching explains the spread. The five-centre study identified it and no trial has randomised on it, which would be a straightforward and useful design.

Telling a real spread from a definitional one

This corpus has spent several territories finding that conflicting figures indicate a definitional problem. Here they mostly do not, and the distinction is worth making operational because getting it wrong runs in both directions.

Three tests separate the cases.

Do the units differ, and does converting them reconcile the figures? 72.6 seconds per encounter and 30 minutes per day are different units. An emergency physician seeing twenty patients saves about 24 minutes, which lands near the ambulatory figure, and the apparent conflict dissolves into arithmetic. A definitional conflict does not dissolve this way: 5% and 19% data readiness stay 5% and 19% however you convert them, because they count different populations.

Does the variation track something in the world that would plausibly cause it? Emergency and ambulatory documentation genuinely differ in length, structure and interruption. A mechanism exists and it predicts the direction. Where a spread has no plausible physical cause, as with a zero-click rate appearing at 22.4% and 60% from one provider, the cause is definitional by elimination.

And do the authors state their measurement? The studies here define time-in-note, encounter and shift explicitly, because clinical journals require it. The spread survives that clarity, which is the strongest single indicator that it is real: definitional conflicts usually collapse the moment both parties state their definitions.

The practical consequence differs sharply. A definitional spread is resolved by agreeing a definition, which is cheap. A genuine spread is not resolved at all, and the correct response is to stop looking for the number and start asking which setting you are in.

And the failure mode this corpus should watch is its own. Having found definitional problems repeatedly, the temptation is to diagnose one everywhere. A method that finds the same fault in every subject has stopped discriminating, and the honest test is whether it can identify a case where the disagreement is simply the world. This is that case, and stating it is more useful than another debunk.

What a health system can check before buying

Unusually for this corpus, the buyer questions here are answerable from published trials rather than from a vendor.

Which setting produced the number you were shown? An emergency department result of 72.6 seconds per encounter and an ambulatory result of 30 minutes per day are both real and neither transfers. A vendor quoting the larger figure to an emergency department is quoting somebody else's setting.

What was the design? Randomised, cohort, or a pre-post survey with no control. The strongest burnout figure in this literature comes from the weakest design, and the strongest design reports the most modest effects.

What is the review policy for one note in fourteen? A fabrication rate is only a defect if nothing catches it. A system that mandates review, samples the reviews, and measures the edit rate has converted a defect into a controlled cost. One that assumes review is happening has not.

And who is coached? The five-centre study found the largest benefits going to clinicians who were trained on the tool rather than left to work it out. That is the only variable in this literature a buyer directly controls, and it is not on any pricing page.

None of these requires a pilot. They are questions about published evidence, answerable before a contract, which is a position no other subject in this corpus has offered.

The counter-argument

The strongest evidence is on the weakest endpoint. The randomised trial's burnout finding is a secondary endpoint its own authors say needs confirmation, and the largest burnout effect comes from a 30-day pre-post survey with no control group. The design quality and the finding strength run in opposite directions here, which is exactly what the corpus warns about elsewhere.

Novelty is not controlled for. Giving overworked physicians a tool that visibly removes a hated task, and surveying them thirty days later, measures relief and enthusiasm alongside any durable effect. A control group receiving a different visible intervention would separate them and none did.

Recruitment was not neutral. Participants were recruited by departmental email and leadership nomination in a system trialling the product, which selects for clinicians disposed to engage. That is standard for pragmatic trials and it bounds generalisation.

And the fabrication figure deserves more weight than this article gives it. One note in fourteen containing invented content, in a legal document that a physician signs, is a serious defect. Treating it as a known limitation to design around, as the vendor-facing literature does, assumes a review process that the same literature says is inconsistently performed.

The short version

Ambient scribes are the first subject in this corpus with registered trials rather than vendor surveys. The UCLA study was a three-arm pragmatic randomised trial, developed under SPIRIT-AI, registered as NCT06792890, reported under CONSORT-AI, comparing two commercial scribes against usual care, and it discloses that institutional urgency precluded timely pre-registration.

It found roughly 7% improvement in burnout scores and remarkably similar performance across both vendors, which is a finding no vendor comparison could have produced.

Time savings range by a factor of twenty-five and the range is real. 72.6 seconds per encounter across 10,336 emergency department encounters, about 30 minutes per provider per day at UW Health, and 16 minutes per eight hours across roughly 1,800 clinicians at five centres. Different settings, different units, and the largest study reports the smallest effect.

Burnout moved more than time explains. A six-system study found the proportion burned out falling from 51.9% to 38.8% in thirty days, an odds ratio of 0.26. Sixteen minutes does not produce a thirteen-point drop, which suggests the benefit is the removal of a dreaded task rather than the recovery of minutes, and that time-in-note was a weaker proxy than assumed.

And roughly one note in fourteen contains fabricated content. Physical examination documentation is unreliable, review is mandatory, review time partially offsets the savings, and no vendor accepts clinical liability. The efficiency is shared and the exposure sits with whoever signs.

Common questions

What makes this evidence different from most AI deployment evidence? It comes from registered trials rather than vendor surveys. The UCLA study was a parallel three-arm pragmatic randomised trial with covariate-constrained randomisation balancing on baseline time-in-note, burnout score and clinic days per week, developed in accordance with SPIRIT-AI, registered on ClinicalTrials.gov as NCT06792890, and reported under CONSORT-AI. It also discloses its own limitation: institutional urgency to test during a brief contractual window precluded timely pre-registration, with the protocol submitted in December 2024 and published in January 2025.

How much time do ambient scribes actually save? It depends heavily on setting, and the range is genuine. A Stanford emergency department retrospective cohort of 10,336 encounters found a 72.6-second reduction per encounter, with a 95% confidence interval of 63.8 to 81.4 seconds. A randomised trial at UW Health reported about 30 minutes less documentation per provider per day. A study of roughly 1,800 clinicians across five academic medical centres, published in early 2026, found about 16 minutes saved per eight hours of patient care and 13 fewer minutes in the record overall. Different units and different settings, with the largest study reporting the smallest effect.

Why did burnout improve more than the time savings explain? Probably because the benefit is not primarily time. A quality improvement study across six health systems found the proportion of clinicians experiencing burnout falling from 51.9% to 38.8% after 30 days, an odds ratio of 0.26, while time savings in the larger studies are measured in minutes. Documentation is not merely long; it is a specific aversive task that follows the working day. Removing a dreaded task is not the same as removing minutes, and time-in-note appears to have been a weaker proxy for the experience than assumed.

How accurate are the notes? Roughly 7% of notes contain fabricated content, approximately one in fourteen. Physical examination documentation is reported as unreliable, which is unsurprising since the system hears the consultation and cannot observe the examination. Physician review is mandatory rather than optional, and the time required for review partially offsets the time saved by ambient capture. Related research is now examining clinician edits to AI drafts and modifications to hedging language, which measures what the reviewer had to fix.

Who carries the risk? The clinician. The system drafts the note, the physician signs it, the signature carries the legal weight, and no vendor accepts clinical liability. The efficiency accrues to the workflow and the exposure stays with the person who signs, which is a cost distribution the time-savings measurements do not capture.

Did the trials find any difference between vendors? No, and this is one of the more useful results. The randomised trial comparing two distinct commercial platforms found remarkably similar performance and reception across both. No vendor comparison or bake-off could have established that, and it removes a procurement question organisations would otherwise spend months on.

What is the strongest objection to the positive findings? That the design quality and the finding strength run in opposite directions. The randomised trial's burnout result is a secondary endpoint its own authors say needs confirmation in larger multicentre trials, while the largest burnout effect comes from a 30-day pre-post survey with no control group, a design that captures novelty and relief alongside any durable benefit. Recruitment was also by departmental email and leadership nomination in systems trialling the product, which selects for clinicians disposed to engage.

Does better evidence produce a single answer? Not here, and that is worth noting rather than treating as a failure. Registered trials produced defensible statements about direction, vendor equivalence and failure rate. They did not produce one effect size, because the effect genuinely varies between an emergency department and an ambulatory clinic. This corpus has spent several territories arguing that conflicting figures usually indicate a definitional problem; sometimes a spread is simply the world, and distinguishing the two cases is the harder skill.

Sources

Primary documents only. Where a claim rests on a single report, the entry says so.

  1. A Randomized-Clinical Trial of Two Ambient Artificial Intelligence Scribes Lukac, Turner, Vangala et al., NEJM AI 2025, registered NCT06792890 The three-arm pragmatic randomised trial: covariate-constrained randomisation on time-in-note, burnout score and clinic days, SPIRIT-AI protocol, CONSORT-AI reporting, vendor equivalence, and the disclosed late pre-registration.
  2. Ambient AI Scribes and Emergency Department Documentation Burden Preiksaitis, Alvarez, Winkel et al., Stanford, July 2026 The retrospective cohort of 10,336 encounters finding a 72.6-second reduction in on-shift documentation time per encounter, 95% CI 63.8 to 81.4, using mixed-effects models with physician random intercepts.
  3. Use of Ambient AI Scribes to Reduce Administrative Burden and Professional Burnout JAMA Network Open, quality improvement study across six health systems The pre-post survey of 131 ambulatory clinicians finding burnout falling from 51.9% to 38.8% after 30 days, odds ratio 0.26. A design with no control group, which is stated in the article.
  4. Ambient AI Scribes in 2026: Clinical Evidence and Vendor Comparison EHR Source The roughly 7% fabrication rate, the unreliability of physical examination documentation, mandatory physician review partially offsetting capture savings, and the absence of vendor clinical liability. A vendor-facing compilation, which is stated.

Further reading

The primary literature behind the claims above, drawn from the concept entries this post links to, so a claim carries the same source here as it does there.

  • Raji et al. (2021), AI and the Everything in the Whole Wide World Benchmark — operationalisation determining what a measurement can support. :: https://arxiv.org/abs/2111.15366 Scope Boundary
  • Wong et al. (2021), External Validation of a Widely Implemented Proprietary Sepsis Prediction Model — the difference between overall performance and performance on the population that matters. :: https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2781307 Scope Boundary
  • Stenberg (2026), The end of the curl bug-bounty — evaluation burden externalised onto an unpaid maintainer until the programme closed. :: https://daniel.haxx.se/blog/2026/01/26/the-end-of-the-curl-bug-bounty/ Cost Externality

Learn the concepts

← All posts