Home/Blog/Evaluation & evidence/The framework that exists produced 1.6%

The framework that exists produced 1.6%

The FDA has two AI tracks. One is final, has authorised over 1,350 devices, and is the regime under which almost none of them cite a trial. The other missed its own deadline five weeks ago.

TL;DR. The FDA runs two separate AI tracks with different scopes, terminology and timelines. The device track has final guidance, including a December 2024 document on predetermined change control plans, and over 1,350 AI-enabled devices authorised by early 2026, roughly double the 2022 count. The drug track has a draft. Published January 2025, informed by CDER's review of more than 500 submissions containing an AI component between 2016 and 2023, it proposes a risk-based seven-step credibility framework tied to a specific context of use, and it consists of non-binding recommendations. Final guidance was signalled for Q2 2026, which ended on 30 June. It remains in draft, with commentary suggesting late 2026 or 2027. And the corpus's own earlier finding sits under the finished track: of 1,524 cleared AI medical devices, 1.6% cite clinical trial data.

---

Status: primary documents are public and the timeline is checkable. Sources are the FDA's own guidance pages and documents, a peer-reviewed critical review in the Journal of Chemistry, and legal and industry analyses. This corpus does not take positions on contested political questions, and whether this is the right regulatory approach is one.

---

Two tracks, not one framework

The device side and the drug side are separate, and conflating them is the commonest error in coverage of this subject.

The device track sits with CDRH and has developed incrementally since 2019: a discussion paper proposing a framework for modifications to AI and machine learning software as a medical device, a 2021 action plan setting out a total product lifecycle approach, and final guidance in December 2024 on predetermined change control plans, which govern how a device designed to be updated after clearance may change without a new submission.

The drug track sits across CDER, CBER and others, and produced its first cross-centre draft in January 2025.

The FDA says the centres coordinate. As of writing they have published separate guidance documents with separate scopes, separate terminology and separate timelines, rather than one rulebook.

Which matters because "the FDA's AI framework" is usually quoted as a single thing, and a claim about one track is routinely used to characterise the other.

The draft, and what it proposes

"Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products", January 2025, developed jointly across CDER, CBER, CDRH, the Center for Veterinary Medicine, the Oncology Center of Excellence, the Office of Combination Products and the Office of Inspection and Investigations.

It was informed by CDER's own experience reviewing more than 500 submissions containing an AI component between 2016 and 2023, which is a substantial evidentiary base for a first guidance document and is the kind of institutional record this corpus has found missing almost everywhere else.

Its core proposal is a risk-based seven-step credibility framework, establishing and documenting the credibility of a model for a specific proposed context of use.

That last phrase is the part worth crediting. Credibility is not established for a model in general. It is established for a model doing a particular job, which is scope boundary written into a regulatory instrument, and it is the correct structure.

And the guidance is non-binding. It provides recommendations on how to demonstrate credibility rather than requirements, which is standard for FDA guidance and is worth stating plainly because the word "framework" implies more.

The deadline that passed

The FDA signalled that final guidance was expected in Q2 2026.

Q2 2026 ended on 30 June. As of the beginning of August it remains in draft, and industry commentary suggests finalisation is unlikely before late 2026 or 2027.

Sponsors are advised to verify status directly with the agency rather than rely on any estimate, which is itself informative about how settled the position is.

This is the same shape the EU AI Act article documented five weeks ago, where high-risk obligations were deferred from 2 August 2026 to December 2027 because the harmonised technical standards required to demonstrate conformity had not been delivered.

Two jurisdictions, two AI regimes, and the same pattern: the obligations requiring an evaluative apparatus slip, and the parts requiring only a decision do not.

The FDA's device track has final guidance because change control is a procedural question. The drug track is still drafting because what counts as credible evidence from a model is a scientific question that nobody has settled.

What the finished track produced

This is the part that bears on whether frameworks fix evidence problems, and the answer is uncomfortable.

By early 2026 the FDA had authorised over 1,350 AI-enabled devices, roughly double the 2022 figure.

This corpus's earlier finding is that of 1,524 cleared AI medical devices, 1.6% cite clinical trial data.

Those two facts describe the same regime. The device pathway has final guidance, a lifecycle approach, change control plans and a decade of incremental development, and it clears devices overwhelmingly on substantial equivalence to existing products rather than on trial evidence.

That is not a failure of the framework. Substantial equivalence is how the pathway is designed to work and has been since long before AI. The framework governs how a device may be modified after clearance, not what evidence is required to clear it.

Which is precisely the point. A reader hearing that the FDA has a comprehensive AI framework will reasonably infer that cleared devices have been evaluated for clinical benefit. The framework is real, it is well constructed, and it does not do that.

The gap the critical review names

*A twenty-page critical review published in the Journal of Chemistry in 2026 credits the structured risk-based credibility framework as a strength and identifies areas needing refinement.*

The gap most often named is generative models.

The draft's treatment of novel model classes is limited, and the questions the final guidance is expected to address include what constitutes a sufficient model risk assessment, how to handle ensemble models with evolving architectures, how to manage AI supplied as a service by external vendors, and whether generative or LLM-based tools require a different framework entirely.

The structural difficulty is straightforward. A credibility framework built around establishing that a model performs reliably for a defined context of use assumes the model's behaviour is stable enough to characterise.

A predictive model trained for one task has that property. A general-purpose generative model does not, and the framework's central mechanism does not obviously extend to it.

That is not an oversight. It is a genuine open problem, it is why the guidance is late, and no regulator anywhere has solved it.

The international position

On 14 January 2026 the FDA and EMA jointly released "Guiding Principles of Good AI Practice in Drug Development", ten high-level principles covering the product lifecycle: human-centric design, a risk-based approach, adherence to standards, clear context of use, multidisciplinary expertise, data governance and documentation, model design and development practices, risk-based performance assessment, lifecycle management, and clear essential information.

Joint publication by two major regulators is a real development and it reduces the divergence problem that faces any sponsor filing in both jurisdictions.

Ten high-level principles are also not a compliance route. They are the layer above the guidance that has not been finalised, and a principle such as "clear context of use" tells a sponsor what to establish rather than how much evidence establishes it.

Which is where the disclosure question actually sits. Principles are cheap to agree and specifications are expensive, and every regime examined in this corpus has produced the first faster than the second.

The two tracks, side by side

Setting them out removes the commonest confusion in this subject.

Device trackDrug and biologics track
CentreCDRHCDER, CBER and five others
StatusFinal guidance, December 2024Draft, January 2025
GovernsPost-clearance modificationCredibility of model evidence
MechanismPredetermined change control planSeven-step credibility framework
Output so far1,350-plus devices authorisedNo finalised route
BindingGuidance, on a statutory pathwayNon-binding recommendations

The fifth row is the one that gets quoted and the third row is the one that matters.

A framework governing how a device may change after clearance is not a framework governing whether it should have been cleared, and the two are constantly merged into "the FDA regulates AI devices."

And the asymmetry in status has a clean explanation. Change control is a procedural question: what may be altered, within what bounds, with what monitoring. It can be answered by drawing a line.

Credibility of model-derived evidence is a scientific question. How much validation makes a model's output usable in a regulatory decision has no obvious line, and drawing one prematurely would be worse than the delay.

Which reframes the lateness. The device track finished first because its question was easier, not because its centre works faster.

What a sponsor can actually do now

The draft is non-binding and unfinished, and it is still the operative document, which puts sponsors in an unusual position.

Establish context of use before anything else. The framework is built around it, and a model characterised for a general purpose cannot be assessed under a structure that assesses fitness for a specific one. This is the step that determines every subsequent one and is the one most often skipped.

Document the risk assessment as a first-class artefact. The seven steps are a documentation discipline as much as an evaluation one, and a sponsor who evaluated well and recorded thinly has the same submission as one who did neither.

Expect the generative question to be unanswered. If a submission depends on a general-purpose model, no settled route exists, and the honest position is early engagement with the agency rather than a framework applied by analogy.

And do not read the device track's maturity as covering this. They are different centres, different pathways and different questions, and the December 2024 final guidance addresses none of it.

None of that is advice, and this corpus is not a regulatory consultancy. It is a description of what the public documents say, offered because the documents are public and most coverage of them is not this specific.

Three things this establishes

"The FDA's AI framework" is two frameworks and the finished one governs something narrower than its reputation. Predetermined change control plans are final and well designed, over 1,350 devices are authorised, and clearance runs on substantial equivalence, which is why 1.6% cite a trial.

Context of use is the right idea and it is the reason the guidance is late. Establishing credibility for a specific job rather than in general is correct, and it presumes behaviour stable enough to characterise, which general-purpose generative models do not have.

And two jurisdictions produced the same pattern within five weeks. The EU deferred what needed harmonised standards; the FDA has not finalised what needs a settled evidentiary standard. The parts requiring an apparatus slip, and principles arrive on time.

What it does not establish

That the FDA is failing. A first cross-centre guidance informed by 500 submissions, delivered within two years of the technology becoming general, is not slow by regulatory standards.

That 1.6% is the framework's fault. Substantial equivalence predates AI by decades and is a statutory pathway rather than a guidance choice.

That the delay is avoidable. How to establish the credibility of a general-purpose generative model for a regulatory decision is an unsolved problem, not a drafting backlog.

And nothing about whether this is good regulation. That is a contested political question and this corpus does not take positions on those.

What is unresolved

When the drug guidance finalises. Signalled for Q2 2026, still draft, with estimates running to 2027.

Whether generative models get a separate framework. Named as the key gap by the critical review, and unaddressed in the draft.

What the device track will require of foundation-model devices. The FDA is reported to be exploring how to identify devices using foundation models and to update its public AI-enabled device list.

And whether any of it changes the 1.6%. The credibility framework governs evidence submitted in support of a decision. Nothing in it changes which decisions require what evidence, which is where the number comes from.

What Territory 11 has now shown about regulation

This territory was chosen because its evidence is good, and the regulatory article is the one that explains why.

Medicine has the strongest evidence infrastructure of any field this corpus has examined. Registered trials, protocol pre-specification, blinding, ethics review, journals that publish nulls, and a regulator with statutory authority.

And each of the five articles found a failure that infrastructure did not prevent.

Scribes: registered trials, and the effect varied twenty-five-fold by setting because the trials were run in different settings and nobody is responsible for the synthesis.

Mammography: three Lancet-family papers, and the claim outran the finding through press releases and coverage, which no journal governs.

Drug discovery: every study registered, and the field-level success rate disputed by twenty-eight points because no register of the category exists.

Diagnosis: randomised trials on a benchmark inherited from medical education, which isolated the part humans find hard.

Therapy chatbots: one trial, and its control group received nothing.

Five failures, and the regulator addresses none of them, which is the finding this article adds rather than a criticism of the FDA.

Because regulation is unit-scoped by construction. A regulator assesses a submission. A submission is one product, for one use, by one sponsor, and every mechanism the FDA has built operates at that level: credibility for a context of use, change control for a device, evidence for a decision.

The failures in this territory are all above that level. Synthesis across settings. Transmission after publication. Aggregation across a category. Choice of benchmark. Choice of comparator.

None of those is a submission and none of them has a regulator.

Which is the aggregate evidence gap with a stronger version of the claim. It is not merely that nobody owns the sum. It is that the most powerful quality mechanism in the field is structurally incapable of owning it, because its authority attaches to products and the problems attach to literatures.

And that generalises past medicine. Any regulator of AI, in any sector, will assess deployments. The failures this corpus has documented across eleven territories are overwhelmingly about how evidence is compared, transmitted and aggregated, and a regulator is the wrong instrument for all three.

The counter-argument

Judging a framework by the 1.6% figure is a category error. Substantial equivalence is a statutory pathway created by Congress, guidance cannot override it, and criticising the FDA's AI work for not fixing 510(k) holds a guidance document responsible for a law. This article states that and then structures itself around the juxtaposition anyway.

Non-binding guidance is more useful than it sounds. Sponsors follow it because deviating invites questions, so recommendations function as requirements in practice, and describing the guidance as merely advisory understates its operational force.

The two-track split is sensible, not a failure of coordination. Drugs and devices have different statutory bases, different review processes and different risk profiles, and demanding one unified rulebook would produce a worse document than two fitted ones.

And the deadline complaint is thin. Regulatory guidance routinely slips, Q2 2026 was a signal rather than a commitment, and a five-week overrun on a document addressing an unsolved scientific question is not evidence of anything. This corpus has criticised others for reading small delays as significant.

The short version

The FDA runs two AI tracks. The device track, with CDRH, has final guidance including a December 2024 document on predetermined change control plans, and over 1,350 AI-enabled devices authorised by early 2026, roughly double 2022.

The drug track has a draft, published January 2025 across seven FDA offices, informed by more than 500 submissions containing an AI component reviewed between 2016 and 2023, proposing a risk-based seven-step credibility framework for a specific context of use, in non-binding form.

Final guidance was signalled for Q2 2026, which ended on 30 June. Still draft, with estimates running to late 2026 or 2027, and sponsors advised to check status with the agency directly.

Which is the EU AI Act pattern again, five weeks apart. What requires an evaluative apparatus slips; what requires only a decision does not. The FDA finalised change control because it is procedural, and has not finalised credibility because what counts as evidence from a model is unsettled.

And the finished track produced the corpus's own earlier finding. Of 1,524 cleared AI medical devices, 1.6% cite clinical trial data, because clearance runs on substantial equivalence and the framework governs post-clearance modification rather than pre-clearance evidence. The framework is real, well constructed, and does not do the thing its reputation implies.

The named gap is generative models. A credibility framework presumes behaviour stable enough to characterise for a defined use. A general-purpose model does not have that property, which is an open problem rather than an oversight, and no regulator anywhere has solved it.

Common questions

Does the FDA have an AI framework? It has two, with different scopes, terminology and timelines. The device track, run by CDRH, has developed since a 2019 discussion paper through a 2021 action plan to final guidance in December 2024 on predetermined change control plans, which govern how a device designed to be updated may change after clearance. The drug and biologics track produced its first cross-centre draft guidance in January 2025 and it remains in draft. The agency says the centres coordinate; they have published separate documents rather than one rulebook.

What does the draft guidance propose? A risk-based seven-step credibility framework for establishing and documenting that an AI model is credible for a specific proposed context of use. It was developed jointly across CDER, CBER, CDRH, the Center for Veterinary Medicine, the Oncology Center of Excellence, the Office of Combination Products and the Office of Inspection and Investigations, and was informed by CDER's review of more than 500 submissions containing an AI component between 2016 and 2023. It provides non-binding recommendations rather than requirements.

Why is context of use the important part? Because credibility is established for a model doing a particular job rather than for a model in general, which is the correct structure and is a scope principle written into a regulatory instrument. It is also the reason the guidance is difficult to finalise for generative systems: establishing credibility for a defined use presumes the model's behaviour is stable enough to characterise, which a predictive model has and a general-purpose generative model does not.

What happened to the deadline? The FDA signalled final guidance for Q2 2026, which ended on 30 June. As of early August it remains in draft, and industry commentary suggests finalisation is unlikely before late 2026 or 2027, with sponsors advised to verify status directly with the agency. That is the same pattern the EU produced five weeks earlier, where high-risk obligations were deferred because the harmonised standards needed to demonstrate conformity had not been delivered.

If the device framework is final, why do so few devices cite trials? Because the framework governs a different question. Predetermined change control plans address how an AI-enabled device may be modified after clearance without a new submission. What evidence is required to clear a device in the first place is governed by the statutory pathway, which for most devices is substantial equivalence to an existing product. That is why of 1,524 cleared AI medical devices, 1.6% cite clinical trial data, and it is a feature of the pathway rather than a failure of the AI guidance.

What is the main gap in the draft? Generative and LLM-based models. A critical review published in the Journal of Chemistry in 2026 credits the structured risk-based credibility framework and identifies the limited treatment of novel model classes as a key gap. Open questions include what constitutes a sufficient model risk assessment, how to handle ensemble models with evolving architectures, how to manage AI supplied as a service by external vendors, and whether generative tools need a different framework entirely.

What is the international position? On 14 January 2026 the FDA and EMA jointly released Guiding Principles of Good AI Practice in Drug Development, ten high-level principles covering human-centric design, a risk-based approach, adherence to standards, clear context of use, multidisciplinary expertise, data governance and documentation, model design and development practices, risk-based performance assessment, lifecycle management and clear essential information. Joint publication reduces divergence for sponsors filing in both jurisdictions, and ten principles are not a compliance route: they sit above the guidance that has not been finalised.

What is the strongest objection to this article? That judging an AI guidance document by the 1.6% figure is a category error, since substantial equivalence is a statutory pathway created by Congress that guidance cannot override. A second objection is that the deadline complaint is thin: regulatory guidance routinely slips, Q2 2026 was a signal rather than a commitment, and a five-week overrun on a document addressing an unsolved scientific question is not evidence of much, which is a standard this corpus has applied to others.

Sources

Primary documents only. Where a claim rests on a single report, the entry says so.

  1. Considerations for the Use of Artificial Intelligence to Support Regulatory Decision Making for Drug and Biological Products FDA draft guidance, January 2025 The primary document: the risk-based credibility framework, context of use, and the reference to CDER's landscape analysis of AI in regulatory submissions.
  2. Artificial Intelligence for Drug Development FDA, CDER The agency's own document list, including the January 2026 joint FDA and EMA Guiding Principles of Good AI Practice in Drug Development.
  3. FDA AI Guidance: Drugs and Medical Devices CASRAI The two-track structure with separate scopes, terminology and timelines, the device track history from the 2019 discussion paper to the December 2024 final change control guidance, and the note that the drug draft remains unfinalised with estimates running to 2027.
  4. A Critical Review of the FDA's Draft Guidance on Artificial Intelligence in Drug and Biological Product Regulation Niazi, Journal of Chemistry, 2026 The peer-reviewed assessment crediting the structured risk-based credibility framework and identifying the limited treatment of generative and novel model classes as the key gap.

Further reading

The primary literature behind the claims above, drawn from the concept entries this post links to, so a claim carries the same source here as it does there.

  • Raji et al. (2021), AI and the Everything in the Whole Wide World Benchmark — operationalisation determining what a measurement can support. :: https://arxiv.org/abs/2111.15366 Scope Boundary
  • Wong et al. (2021), External Validation of a Widely Implemented Proprietary Sepsis Prediction Model — the difference between overall performance and performance on the population that matters. :: https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2781307 Scope Boundary

Learn the concepts

← All posts