Home/Blog/Safety & governance/AI in medicine: 1,524 devices, 1.6% with trial data
AI in medicine: 1,524 devices, 1.6% with trial dataOf 691 devices reviewed, 1.6% cited randomised trial data100%FDA-cleared devices2%cited a randomised trialOF ORGANISATIONS DEPLOYING AIOf 691 devices reviewed, 1.6% cited randomised trial data
Of 691 devices reviewed, 1.6% cited randomised trial data

AI in medicine: 1,524 devices, 1.6% with trial data

The FDA lists 1,524 AI-enabled medical devices. A review of 691 found 1.6% cited a randomised trial and under 1% reported patient outcomes. Medicare pays for about ten.

TL;DR. Regulatory clearance in medical AI is a claim about similarity to an existing device, not a claim about patient benefit. Of 691 cleared devices reviewed, 1.6% cited randomised trial data and under 1% reported patient health outcomes. About three quarters of all authorisations are radiology. Medicare has assigned payment to roughly ten of the 1,524 devices on the list. What is unambiguously working is documentation: ambient scribes are used by over 200,000 clinicians and save around five minutes per encounter, and they work because a clinician reads every word before it enters the record. What is not working is autonomous diagnosis, and no cleared device claims to do it.

---

The FDA's public catalogue of AI-enabled medical devices listed 1,524 entries as of March 2026.

A cross-sectional review of 691 of those devices found that 1.6% cited data from a randomised clinical trial, and fewer than 1% reported actual patient health outcomes. A separate analysis of 1,016 authorisations found that nearly half of the FDA summaries did not describe the study design used for clearance, and more than half omitted the sample size.

As of mid-2024, the Centers for Medicare and Medicaid Services had assigned payment to around ten of them.

Those four numbers are the article. Clearance in this field is a statement about resemblance to a device already on the market, not a statement about whether patients do better. The gap between the two is the single most important thing to understand about medical AI, it is measurable, and almost nobody states it plainly because almost everybody writing has something to sell.

What clearance actually certifies

The distinction that resolves most of the confusion, and it is regulatory rather than technical.

Nearly every AI medical device reaches the market through the 510(k) pathway. That pathway asks a manufacturer to demonstrate substantial equivalence to a legally marketed predicate device. The question it answers is whether the new thing is sufficiently like the old thing. It is not whether the new thing helps.

This is not a loophole. It is the pathway working as designed, for a category it was designed for, which was physical instruments where equivalence is a reasonable proxy. A new blood pressure cuff that measures the same quantity the same way as a cleared cuff probably works about as well. Software that learns a decision boundary from a training set does not have the same relationship to its predicate.

De Novo is the other route, used where no predicate exists, and it carries a higher evidentiary burden. It is rare. Paige Prostate received the first De Novo clearance for an AI pathology product in 2021, which is notable precisely because it was the first.

The practical consequence: "FDA cleared" tells you a device exists in a regulated category. It does not tell you it was tested against patient outcomes, and in the overwhelming majority of cases it was not.

Where the devices actually are

The distribution is far more concentrated than the coverage suggests.

Radiology accounts for roughly 76 to 80% of all FDA-authorised AI medical devices, and this proportion has held across independent counts from 2021 through 2026. Cardiology is about 10%. Everything else, neurology, pathology, ophthalmology, dentistry, shares the remainder.

The manufacturer concentration is similar. GE HealthCare leads with 120 cumulative authorisations, then Siemens Healthineers at 89, Philips at 50, Canon at 45.

Two things follow from this that are usually missed.

The field is an imaging field. When someone says AI is transforming medicine, the accurate version is that AI is doing measurable work in image analysis and very little elsewhere. That is a real achievement and it is a narrower claim.

The incumbents own it. The top four manufacturers by clearance count are the companies that already sold the scanners. This is not a disruption story. It is an incremental feature being added by the existing equipment vendors, which is what most successful medical technology adoption looks like.

What is unambiguously working

Three categories have real evidence behind them, and it is worth being specific about why.

Ambient documentation

The most widely deployed clinical AI application, and the one with the least hype attached.

Ambient scribe systems listen to a clinician-patient conversation and draft a structured note into the electronic record. Over 200,000 clinicians now use one such system, and reported time savings run to roughly five minutes per patient encounter.

The context that makes that number matter: physicians spend approximately two hours on electronic record documentation for every hour of direct patient care. Documentation burden is among the most-cited drivers of clinician burnout. A tool that returns five minutes an encounter across a full clinic day is doing something real.

And note the design that makes it safe. The system drafts; it does not file. A clinician reviews and signs every note before it enters the record. The model's output is an input to a human decision, and the human is the one who is licensed and accountable.

That is the whole reason this works while more ambitious deployments do not, and it generalises: the successful medical AI deployments are the ones where a model produces a draft that a qualified person checks, and the unsuccessful ones are those where a model produces a decision.

Triage and prioritisation

The second working category, and it is subtler than it sounds.

Triage tools do not diagnose. They reorder a worklist. A system that flags a suspected intracranial haemorrhage moves that study to the top of the radiologist's queue. The radiologist still reads it and still decides.

The value is entirely in latency. If a bleed is read in eleven minutes rather than fifty, the patient benefits, and nothing about the diagnostic process changed. One 2026 clearance covers a single body CT triage solution across fourteen conditions, with reported mean sensitivity of 97% and specificity of 98% in its pivotal study.

Triage is the shape of medical AI that regulators and clinicians are most comfortable with, because a false positive costs an unnecessary early read and a false negative returns the study to the normal queue rather than to nowhere. The failure mode is bounded.

Screening at scale, where access is the constraint

The third category is where the comparator is not an expert but nothing at all.

Diabetic retinopathy screening is the established case. In settings without an ophthalmologist, an autonomous screening system is measured against no screening. That is a far lower bar than measuring against a specialist, and clearing it produces real benefit.

Deployments in this shape are underway at scale, including programmes providing millions of free tuberculosis, lung cancer and breast cancer screenings in India over the coming decade.

The evidentiary standard should change with the comparator, and this is the case where enthusiasm is most justified. A system that catches 80% of cases in a population currently receiving 0% screening is a substantial gain. The same system offered as a replacement for specialist reading in a well-resourced hospital would be a downgrade.

The four bars, and why they are not the same test

Most disagreement about medical AI comes from two people using the word "works" against different comparators. Naming them makes the disagreement resolvable.

Bar one: better than nothing. No screening exists in this setting. A system catching 70% of cases where 0% were previously caught is a large gain, and demanding specialist-equivalence here argues for the status quo, which is nobody being screened. This is the correct bar for retinopathy screening in a district with no ophthalmologist.

Bar two: better than the available non-specialist. A general practitioner reading a scan a radiologist would normally read. Common in practice, rarely stated in marketing, and the bar most triage claims are implicitly measured against.

Bar three: as good as a specialist. The bar most people assume is being claimed and the one least often tested prospectively. Retrospective studies clear it regularly on curated data. Deployment studies clear it far less often.

Bar four: better than a specialist using the tool. The only bar that matters for a well-resourced hospital, because that is the actual alternative. Almost nothing is measured against it, and the results that exist are mixed, because a specialist plus an imperfect tool can be worse than a specialist alone when the tool is wrong confidently.

A claim is not evaluable until you know which bar it cleared. A vendor saying "outperforms clinicians" has usually cleared bar three retrospectively, and a hospital hearing it usually assumes bar four prospectively. Neither party is lying, and they are discussing different things.

The practical form of this is a single question to ask any vendor: compared with what, and measured how? If the answer is "compared with clinicians" without specifying which clinicians under what conditions, the claim has not been made yet.

The sepsis case, which shows the gap most clearly

If you want one example of clearance-versus-benefit, take sepsis prediction. Sepsis kills roughly 270,000 Americans a year and costs more per hospitalisation than any other condition, so the incentive to predict it is enormous.

The most widely deployed model was built into a dominant electronic record system and rolled out across hundreds of hospitals.

A large external validation published in 2021 reported an area under the curve of 0.63. The vendor's later version reports 0.82 to 0.92. Both figures circulate, they measure different model versions under different conditions, and the gap between an internally reported 0.9 and an externally validated 0.63 is the entire problem in one comparison.

The number that matters more than either is the number needed to treat: 21 to 35 patients flagged to catch one event. That is not a scandal on its own; screening tools routinely have unfavourable ratios. It is the operational consequence that matters, and clinicians describe it directly as heavy alert fatigue.

An alert that fires on twenty-one patients to find one real case trains the people receiving it to dismiss it. A prediction nobody acts on has a clinical utility of zero regardless of its AUC, and this is the failure mode that no accuracy metric captures.

The regulatory picture moved in 2024, when the first machine-learning sepsis device received FDA authorisation. Its design is instructive: it is intended to be used alongside an existing clinical suspicion of infection, such as when a blood culture has been ordered, rather than as an ambient screen running on every patient. The narrower deployment is what made the evidence tractable.

The reimbursement signal, and why it is the honest one

Roughly ten of 1,524 devices have assigned Medicare payment.

Reimbursement is a harder gate than clearance because it asks a different question. Clearance asks whether the device is like an existing device. Payment asks whether it produces enough value to be worth public money, and the body deciding is not the body that cleared it and has no interest in the manufacturer's success.

So the reimbursement rate is the closest thing to an independent verdict on medical AI that currently exists, and the verdict is that under 1% has cleared it.

This is worth weighing carefully in both directions. Reimbursement lags clearance by years, so a low rate partly reflects timing. Some truly valuable tools are paid for inside existing bundled payments and never get a separate code. And CMS is conservative by design.

But the same argument was available five years ago, and the count has not moved much. At some point a persistent gap stops being lag and starts being an answer.

What the evidence base actually looks like

Four findings, and the fourth is the one that should change how you read any claim in this field.

Performance in curated studies exceeds performance in deployment, consistently. Equipment variation, differences in patient populations and workflow integration all degrade real-world results relative to retrospective evaluation on clean data.

The reporting is thin. Nearly half of FDA summaries do not describe the study design. Over half omit the sample size. A clinician trying to judge whether a tool will work in their population frequently cannot, because the information required is not in the public record.

Recalls happen. 5.8% of reviewed devices were eventually recalled, mostly for software defects rather than for clinical harm. That rate is neither alarming nor negligible, and it is a reminder that these are software products subject to software failure modes.

And the strongest results are on benchmarks, not patients. A model scoring 91.1% on a medical licensing-exam question set is answering multiple-choice questions written to have one correct answer. That is a meaningful capability measurement and it is not evidence of clinical benefit, because the clinical task is not multiple choice, the differential is not supplied, and the patient in front of you has not been written as a vignette.

How to read a medical AI claim

Six questions, and they are the same six that apply to any AI claim with the medical specifics filled in.

Was it cleared through 510(k) or De Novo? The first says it resembles something already sold. Ask what the predicate was.

Against what comparator was it measured? Against a specialist, against a non-specialist, against a resident, or against nothing? Each is a different claim, and the last is the easiest bar to clear.

Was the validation internal or external? A model validated only on data from the institution that built it tells you almost nothing about your institution. The sepsis case is the standing example: 0.9 internally, 0.63 externally.

What is the number needed to treat, or the alert rate? Sensitivity and specificity do not tell you how many alerts a clinician receives per real event, and that number determines whether the tool gets used or ignored.

Does it report a patient outcome, or a diagnostic metric? Under 1% of cleared devices report the former. If a claim is about accuracy rather than about patients getting better, say so precisely.

And who reviews the output before it acts? The deployments that work put a licensed human between the model and the record. Any claim of autonomy should be read against the fact that essentially no cleared device makes one.

What is unresolved

Whether the 510(k) pathway is appropriate for adaptive software. A device that learns after deployment is not the same device that was cleared, and the regulatory framework for handling that is still being constructed. The current answer is largely to prohibit post-market learning, which is a workaround rather than a solution.

Whether the evidence gap is a transition or a steady state. The optimistic reading is that trials take years and the RCT count will rise. The pessimistic reading is that trials are expensive, the market rewards clearance rather than evidence, and nobody has an incentive to run the study that might return a null result. Both are consistent with the current data.

How to validate on populations the training data underrepresented. Demographic gaps in medical datasets are documented and persistent. Synthetic data is proposed as a mitigation and the validation of synthetically-trained models against real outcomes remains an open research question rather than an established practice.

What happens to clinician skill. If a triage system reliably surfaces the urgent study, the skill of finding it unaided gets exercised less. Whether that matters, over what timescale, and what happens when the system is unavailable are questions nobody has good data on.

And whether generalist models change the analysis. The regulatory framework was built for devices with a defined intended use. A general model answering arbitrary clinical questions does not have one, and it is currently unclear whether it is a device at all.

The counter-argument

The evidence bar being demanded here is higher than medicine applies to itself. A large fraction of accepted clinical practice has never been through a randomised trial, surgical techniques are adopted on far less evidence than 1.6%, and holding software to a standard the rest of the field does not meet is not obviously consistent. The honest version of the complaint is that medical AI should be held to the standard of comparable medical software, not to the standard of a new drug.

Clearance was never supposed to certify benefit, and criticising it for not doing so misreads the system. The FDA regulates safety and effectiveness for an intended use. Clinical value is meant to be established afterwards, by the literature and by payers. The reimbursement gate is the system working, not failing.

And the working deployments are truly valuable. Five minutes per encounter across 200,000 clinicians is an enormous aggregate return, and it does not require a randomised trial to be worth having. An article that leads with 1.6% risks implying that nothing works, when the accurate statement is that the things that work are narrower and less glamorous than the coverage suggests.

The screening case deserves particular care. In a setting with no ophthalmologist, an imperfect autonomous system is not competing with excellence. Applying a rich-country evidentiary standard to a deployment whose comparator is nothing is a way of arguing for the status quo, and the status quo there is no screening at all.

The short version

The FDA lists 1,524 AI-enabled medical devices. A review of 691 found 1.6% cited randomised trial data and under 1% reported patient health outcomes. An analysis of 1,016 found nearly half of the summaries omitted the study design and over half omitted the sample size. Medicare has assigned payment to roughly ten.

Clearance certifies resemblance, not benefit. Nearly all of these devices enter through 510(k), which asks whether the product is substantially equivalent to something already marketed. That pathway was designed for physical instruments where equivalence is a fair proxy, and software that learns a decision boundary does not have that relationship to its predicate.

Roughly 76 to 80% of authorisations are radiology, and the leading manufacturers by clearance count are the companies that already sold the scanners. This is an imaging field, and an incumbent one.

Three things work. Ambient documentation, used by over 200,000 clinicians, saving about five minutes an encounter against a baseline of two hours of records work per hour of patient contact. Triage, which reorders a worklist without diagnosing anything, where the failure mode is bounded. And screening where the comparator is no screening at all rather than a specialist.

What they share is that a licensed human reads the output before it does anything. The deployments that fail are the ones where a model produces a decision instead of a draft.

The sepsis case shows the gap most clearly: an externally validated AUC of 0.63 against internally reported figures above 0.9, and a number needed to treat of 21 to 35, producing what clinicians describe as heavy alert fatigue. A prediction nobody acts on has a clinical utility of zero whatever its AUC, and no accuracy metric captures that.

The reimbursement rate is the honest signal. Clearance asks whether a device resembles an existing one, and the body that decides has no stake in the answer being no. Payment asks whether it is worth public money, and the body that decides has every reason to say no. Under 1% has passed the second gate, and that gap has not closed in five years.

Common questions

Is AI actually used in hospitals today? Yes, extensively, and mostly in ways that get little attention. The most widely deployed application is ambient clinical documentation, where a system drafts a clinical note from a recorded consultation and a clinician reviews and signs it, used by over 200,000 clinicians. Imaging triage, which reorders a radiologist's worklist without diagnosing anything, is the second. Autonomous diagnosis is essentially not deployed, and no cleared device claims to perform it.

How many AI medical devices has the FDA approved? The FDA's public catalogue listed 1,524 AI-enabled devices as of March 2026, though the agency cautions the list is not comprehensive and reflects devices identified through AI-related terminology in authorisation summaries. Roughly 76 to 80% are radiology. The word "approved" is also imprecise: nearly all were cleared through the 510(k) pathway, which is a different and lower bar than approval.

Does FDA clearance mean an AI tool improves patient outcomes? No, and this is the most consequential misunderstanding in the field. Clearance through 510(k) means the device was shown to be substantially equivalent to a legally marketed predicate. A review of 691 cleared devices found only 1.6% cited randomised clinical trial data and fewer than 1% reported actual patient health outcomes. Clearance certifies that a device belongs in a regulated category, not that patients do better with it.

What is the difference between 510(k) and De Novo clearance? 510(k) requires demonstrating substantial equivalence to an existing marketed device, and is how nearly all AI medical devices reach the market. De Novo is used when no suitable predicate exists and carries a higher evidentiary burden. De Novo clearances are rare in AI: the first for a pathology product was granted in 2021, which is notable precisely because it was the first.

Why is most medical AI in radiology? Because imaging is the medical data type that most closely resembles what these systems are good at: a fixed-size input, a large labelled corpus, and a task with a defined answer. It is also where the incumbent equipment manufacturers already sell, which is why the top four companies by clearance count are the companies that already sold the scanners. Roughly three quarters of all authorisations sit in this one specialty.

Can AI diagnose disease better than a doctor? On narrow, well-defined tasks under controlled conditions, some systems match or exceed specialist performance in retrospective studies. Real-world performance consistently falls short of those results because of equipment variation, different patient populations and workflow integration. More importantly, the deployed systems are not designed to diagnose: they are decision support for a licensed clinician, and the regulatory clearances reflect that intended use rather than autonomy.

What went wrong with AI sepsis prediction? The most widely deployed model reported strong internal performance and, in a large external validation published in 2021, an area under the curve of 0.63. The operational problem is worse than the accuracy problem: the number needed to treat runs from 21 to 35 patients flagged per event detected, producing what clinicians describe as heavy alert fatigue. An alert that fires on twenty-one patients to find one case trains its recipients to dismiss it, and a prediction nobody acts on has no clinical utility whatever its AUC.

How should a hospital evaluate a medical AI product? Ask which pathway cleared it and what the predicate was. Ask what comparator it was measured against, since a specialist, a non-specialist and no screening at all are three very different bars. Ask whether validation was internal or external, because the sepsis case shows a gap from above 0.9 to 0.63 between the two. Ask for the alert rate or number needed to treat rather than sensitivity alone. And ask whether any published result reports a patient outcome rather than a diagnostic metric, since under 1% of cleared devices do.

Learn the concepts

← All posts