Home/Blog/Evaluation & evidence/Phase I improved. Phase II did not.

Phase I improved. Phase II did not.

AI-designed drugs clear safety trials at well above industry rates. At the stage that tests whether a drug works, the sources contradict each other, and the more careful ones report no advantage at all.

TL;DR. As of mid-2026 there are between 173 and 200-plus AI-discovered drug programmes in clinical development, roughly 94 in Phase I, 56 in Phase II, 15 in Phase III, and zero regulatory approvals. Phase I success runs at 80 to 90% against a historical industry average near 50 to 65%, which is consistent across sources and is a real result. Phase II is where the sources break. One reports 68% against 30 to 45% traditional. Another reports around 40%. A narrative review covering 2020 to July 2026 states plainly that Phase II success rates align with historical industry averages of roughly 40%. That distinction is the whole subject, because Phase I tests whether a molecule is safe and Phase II tests whether it works. The evidence supports the first claim and does not yet support the second.

---

Status: one consistent finding, one contested one, and no approvals. The rentosertib Phase IIa result is published in Nature Medicine and registered as NCT05938920. The pipeline counts and success rates come from industry analyses that disagree, and are reported as disagreeing rather than averaged. Sources also give two different placebo figures for the same trial, which is noted where it appears.

---

The scoreboard

Roughly 94 programmes in Phase I, 56 in Phase II, 15 in Phase III. Totals are reported as 173 by one count and 200-plus by another, both dated early 2026.

Zero AI-discovered drugs have received regulatory approval.

Analysts put the probability of a first approval at roughly 60%, with the window given variously as 2026 to 2027 or 2027 to 2028 depending on source.

And "AI-discovered" has no standard definition. The label spans everything from a molecule whose target and structure were both generated by models to a programme where machine learning assisted one optimisation step. One analysis notes this explicitly, which is more than most do.

That definitional spread matters for every success rate below, because a category assembled loosely will include programmes that would have looked identical without AI.

The finding that holds

Phase I success for AI-derived molecules runs at 80 to 90%, against a historical industry average around 50 to 65%. One source gives 81% against 52%.

This is reported consistently and it is a genuine result.

Phase I asks whether a compound is safe and tolerable in humans. Failures there are typically toxicity, poor pharmacokinetics, or unacceptable side effects.

Which is exactly what computational design should be good at. Absorption, distribution, metabolism, excretion and toxicity are properties of a molecule's structure, they are the subject of decades of accumulated data, and predicting them is a well-posed problem with a large training set.

One programme illustrates the speed advantage alongside it. Rentosertib reached Phase II 30 months from project inception, against a traditional 4 to 5 years to reach Phase I, at an estimated $50 to $100 million to Phase II.

So the honest version of the AI drug discovery claim is specific and defensible: molecules can be designed faster, cheaper, and with better safety profiles than the historical average.

The finding that does not

Phase II asks whether the drug works in patients with the disease. It is where efficacy is measured for the first time, and it is where the pharmaceutical industry has always lost most of its candidates.

And here the sources contradict each other.

One industry analysis reports Phase II success at 68% against 30 to 45% traditional, which would be transformative.

Another reports around 40% against a traditional 29 to 40%, which is no advantage.

And a narrative review covering 2020 to July 2026 states that Phase II success rates align with historical industry averages of roughly 40%, with the gap between computational promise and demonstrated patient benefit remaining substantial.

A third analysis, the one most often cited as the field's reference point, is Jayatunga and colleagues in Drug Discovery Today, June 2024, titled "How successful are AI-discovered drugs in clinical trials?" It reports the Phase I advantage clearly and Phase II as the point where the picture changes.

Sixty-eight percent and forty percent cannot both describe the same population, and the more cautious sources, including the narrative review and the analysis whose title is the question, report the lower figure.

Why the split makes mechanistic sense

The two stages test different things, and AI's advantages map onto only one of them.

Phase I failure is chemistry. Toxicity and pharmacokinetics are molecular properties, predictable from structure, with decades of data to learn from. A generative model optimising for drug-likeness and ADMET is working on a problem where the answer is in the molecule.

Phase II failure is biology. A drug fails Phase II when the target turns out not to matter for the disease, when the effect is too small in real patients, or when the disease is heterogeneous in ways the model of it did not capture.

None of that is in the molecule. It is in whether the biological hypothesis was right, and target validation is a question about disease mechanism rather than about chemistry.

Which predicts exactly the pattern reported: a large advantage where the problem is molecular, and no advantage where the problem is biological.

And it clarifies what the technology has actually demonstrated. Generative chemistry works. Target selection, which is the harder and more valuable half, has not been shown to work better, and the field's own pipeline data is the evidence.

The most advanced result, read carefully

Rentosertib, also called ISM001-055, is a TNIK inhibitor for idiopathic pulmonary fibrosis from Insilico Medicine. It is described as the first molecule where both the biological target and the chemical structure were identified by proprietary generative models.

*Its Phase IIa was a double-blind, placebo-controlled 12-week study in 71 patients, registered as NCT05938920, published in Nature Medicine on 3 June 2025.*

At 60 mg once daily, mean forced vital capacity improved by 98.4 mL from baseline at 12 weeks, with a 95% confidence interval of 10.9 to 185.9 mL.

Three things about that interval are worth stating.

It is very wide. A point estimate of 98.4 with bounds at 10.9 and 185.9 spans nearly the entire plausible range of clinically meaningful effects.

Its lower bound barely clears zero, which in a 71-patient trial is what a genuine signal looks like at this stage and is not a demonstration.

And the placebo comparison is reported inconsistently across sources, at −20.3 mL in one and −62.3 mL in another. Both describe the same trial. The corpus has not resolved which is correct, and a reader encountering either figure alone would compute a different treatment difference.

None of this diminishes the result. A Phase IIa efficacy signal in IPF, a disease where progression is rarely halted, is a genuine milestone and the field's most significant to date.

It is also, as researchers close to it say, a starting point rather than a finish line, and rentosertib requires a larger Phase III before any regulator considers it.

Where the caution comes from inside the field

Worth quoting because the strongest scepticism is not external.

Insilico's own chief executive has been direct: the algorithm identifies a candidate in months, and getting it into a patient takes years, costs hundreds of millions, and depends on factors no model can predict.

The narrative review's conclusion is similarly plain: the clinical development challenges that define pharmaceutical R&D, target validation, patient heterogeneity, toxicology, pharmacokinetics and trial failure, remain unchanged.

And one analysis names the selection problem directly. Early AI programmes may target easier indications, which would inflate success rates at every stage without the technology contributing anything.

That last point deserves weight. A field choosing tractable targets to demonstrate a platform is behaving rationally, and the resulting success rates measure the choice as much as the method.

The pipeline, stage by stage

Setting the counts against the success rates shows what the field is waiting for.

StageProgrammesAI success rateHistoricalWhat it tests
Phase I~9480 to 90%50 to 65%Safety, tolerability, pharmacokinetics
Phase II~5640% or 68%, disputed29 to 45%Whether it works in patients
Phase III~15No data50 to 60%Whether it works at scale
Approval0none yetnone yetEverything

The rows narrow the way a pipeline should, and the interesting column is the fourth.

The advantage is at the stage testing molecular properties, the dispute is at the stage testing biology, and there is no data at all at the stage that decides.

Fifteen Phase III programmes is the number to watch. It is large enough that outcomes will be informative within a few years and small enough that a single result will move the perceived picture more than it should.

And the zero in the bottom row is doing less work than it appears to. With programmes entering trials from around 2020 and development running ten to fifteen years, zero approvals in 2026 is the expected value rather than a disappointing one, which is the strongest objection to reading the table as a verdict.

What the table does support is narrower and firmer: the field has demonstrated something at Phase I, has not demonstrated it at Phase II, and has produced no evidence at all beyond that.

What would settle it

Three measurements, and only one requires waiting.

Publish the programme list behind each success rate. The 40% and 68% figures for Phase II cannot both be right, neither analysis publishes its denominator, and the disagreement is resolvable by disclosure rather than by research. This is the single cheapest fix available in the subject.

Compare indication difficulty across AI and non-AI cohorts. The concern that early AI programmes target easier diseases is named in the literature, never quantified, and testable with data that already exists. If AI candidates are concentrated in indications with historically higher Phase II rates, the comparison currently being made is invalid, and if they are not, one objection disappears.

And define the term. A standard for what counts as AI-discovered, distinguishing model-generated target and structure from machine-learning-assisted optimisation, would let every subsequent rate mean something. Nothing prevents an industry body from publishing one.

Only Phase III outcomes require time, and the other two are disclosure problems of exactly the kind the previous territory found everywhere: the party holding the data has no obligation to publish and the finding might be uncomfortable.

Which is the recurring shape, arriving in a field with peer review, registered trials and regulators. Good evidence infrastructure at the trial level does not produce good evidence at the field level, because nobody is responsible for the aggregate.

Three things this establishes

The advantage is real and it is at the stage where the problem is molecular. 80 to 90% Phase I success against 50 to 65% is consistent across sources and mechanistically explicable. Generative chemistry does what it claims.

The claim that matters is contested and the careful sources say no. Phase II is reported at both 68% and 40%, and the narrative review covering six years states that rates align with historical averages. The stage that tests whether a drug works has not shown an advantage.

And zero approvals after 173 to 200 programmes is the fact that anchors everything. The pipeline is real, the first approval is plausibly two years away, and until one arrives the entire case rests on intermediate endpoints in a field where intermediate endpoints are famously unreliable.

What it does not establish

That AI drug discovery is failing. A pipeline that did not exist five years ago now has 15 programmes in Phase III, and the speed and cost advantages are substantial and independently reported.

That rentosertib will not work. The Phase IIa signal is real, the disease is one where progression is rarely halted, and the confidence interval excludes zero.

That the 68% figure is wrong. It may describe a different or later cohort, and neither figure comes with a stated denominator this article could check.

And nothing about which analysis is correct on the placebo arm. Two published figures for one trial appear in circulation and this corpus has not resolved them.

What is unresolved

What "AI-discovered" means. No standard definition exists, the label spans wildly different degrees of involvement, and every success rate inherits the ambiguity.

Whether Phase II rates are 40% or 68%. The disagreement is the central factual question in the subject and no source publishes the underlying programme list.

Whether indication selection explains the advantage. Named as a concern, never quantified, and testable by comparing indication difficulty across AI and non-AI cohorts.

And what the first approval will show. A single approval will be read as validation of the field, and one approval from a pipeline of 200 is roughly what chance would produce.

Rigour at the trial, nothing at the field

This is the third subject where every individual unit is rigorously produced and the aggregate is unreliable, and the pattern deserves naming.

Each rentosertib datapoint is registered, blinded, placebo-controlled and peer-reviewed. The trial has a ClinicalTrials.gov identifier, a published protocol and a Nature Medicine paper.

And the field-level Phase II success rate is reported at 40% and at 68% by different analysts, neither of whom publishes a programme list.

Those two facts sit together comfortably because they operate at different levels. Trial rigour is enforced by regulators, ethics committees, journals and registries, all of which govern one study. Nobody governs the sum.

The corpus has seen this twice before.

Medical devices: 1,524 AI devices each cleared through an FDA process, and it took a dedicated study to establish that 1.6% cite clinical trial data. Every clearance was procedurally correct and nobody was counting.

Agent benchmarks: individual benchmarks are carefully constructed and peer-reviewed, and no aggregate of production performance exists because no party is responsible for producing one.

The mechanism is that quality control attaches to the artefact and not to the population of artefacts. A journal reviews a paper. A regulator reviews a submission. An ethics committee reviews a protocol. Each does its job well, and the question "what does all of this add up to" has no owner, no process and no venue.

Which is why the field-level number is where the distortion lives even in domains with excellent per-unit standards, and why a reader encountering a summary statistic about a well-regulated field should ask who computed it and from what.

It also suggests the cheapest available intervention. Not more rigour per trial, which is already high, but a published register of what is in the category and how each entry was counted. For AI drug discovery that is a spreadsheet, and its absence is why two analysts can differ by 28 percentage points on the field's central question.

The counter-argument

Judging the field by approvals is premature by construction. Drug development takes ten to fifteen years, AI-designed molecules began entering trials around 2020, and demanding approvals in 2026 asks for something the timeline cannot yet supply. The absence of approvals is close to uninformative.

The Phase II comparison may be unfair. AI-derived candidates are disproportionately in areas like fibrosis and oncology with historically poor Phase II rates, and a like-for-like comparison by indication does not exist. A raw 40% against a raw 40% may conceal an advantage or a disadvantage.

Speed and cost are real benefits that this framing underweights. Thirty months to Phase II against four to five years to Phase I, at $50 to $100 million, changes the economics of attempting a target at all, which matters most for diseases nobody could previously afford to pursue.

And the molecular-versus-biological split is tidier than the evidence. Target identification in the rentosertib programme was itself model-driven, and if that target proves correct it is direct evidence against the claim that AI has not demonstrated anything about biology. This article's mechanism is a hypothesis that the field's own flagship case may already contradict.

The short version

As of mid-2026: 173 to 200-plus AI-discovered drug programmes in clinical development, roughly 94 in Phase I, 56 in Phase II, 15 in Phase III, and zero approvals.

Phase I success runs at 80 to 90% against a historical 50 to 65%, reported consistently, and mechanistically explicable: Phase I failure is chemistry, and toxicity and pharmacokinetics are molecular properties with decades of training data behind them.

Phase II is contested. One source reports 68% against 30 to 45% traditional. Another reports around 40%. A narrative review covering 2020 to July 2026 says rates align with historical industry averages of roughly 40%, with the gap between computational promise and demonstrated patient benefit remaining substantial.

Phase II failure is biology: the target does not matter, the effect is too small, or the disease is heterogeneous in ways the model missed. None of that is in the molecule, which predicts precisely the pattern reported.

The flagship result is rentosertib, a TNIK inhibitor for IPF, Phase IIa in 71 patients, published in Nature Medicine, showing 98.4 mL FVC improvement at 60 mg with a 95% interval of 10.9 to 185.9 mL. A genuine signal, a wide interval, a lower bound barely clearing zero, and two different placebo figures in circulation for the same trial.

And the strongest caution comes from inside. Insilico's own chief executive notes the algorithm finds a candidate in months while getting it to a patient takes years and depends on factors no model predicts. The review's conclusion is that target validation, patient heterogeneity, toxicology and trial failure remain unchanged.

Common questions

How many AI-discovered drugs have been approved? None. As of mid-2026 there are between 173 and 200-plus AI-discovered drug programmes in clinical development, roughly 94 in Phase I, 56 in Phase II and 15 in Phase III, with zero regulatory approvals. Analysts put the probability of a first approval at roughly 60%, with the window given variously as 2026 to 2027 or 2027 to 2028.

Do AI-designed drugs succeed more often? In Phase I, clearly. Success runs at 80 to 90% against a historical industry average around 50 to 65%, with one source giving 81% against 52%, and the finding is consistent across sources. In Phase II the sources contradict each other: one reports 68% against 30 to 45% traditional, another reports around 40%, and a narrative review covering 2020 to July 2026 states that rates align with historical averages of roughly 40%.

Why would the two stages differ? Because they test different things. Phase I failure is usually chemistry: toxicity, pharmacokinetics and tolerability are molecular properties, predictable from structure, with decades of accumulated data behind them, which is exactly what generative design optimises for. Phase II failure is biology: the target turns out not to matter for the disease, the effect is too small in real patients, or the disease is heterogeneous in ways the model of it did not capture. None of that is in the molecule.

What did the rentosertib trial actually show? A double-blind, placebo-controlled 12-week Phase IIa in 71 patients with idiopathic pulmonary fibrosis, registered as NCT05938920 and published in Nature Medicine in June 2025. At 60 mg once daily, mean forced vital capacity improved by 98.4 mL from baseline at 12 weeks, with a 95% confidence interval of 10.9 to 185.9 mL. That is a genuine efficacy signal in a disease where progression is rarely halted, with a very wide interval whose lower bound barely clears zero in a small trial. Sources also report two different placebo figures for the same study, at −20.3 mL and −62.3 mL.

Is the speed advantage real? Independently reported and substantial. Rentosertib reached Phase II 30 months from project inception, against a traditional four to five years merely to reach Phase I, at an estimated $50 to $100 million to Phase II. That changes the economics of attempting a target at all, which matters most for diseases nobody could previously afford to pursue.

What is the biggest methodological problem? That "AI-discovered" has no standard definition. The label spans molecules where both target and structure were model-generated and programmes where machine learning assisted a single optimisation step, and every success rate in the field inherits that ambiguity. A related concern, named in the literature and never quantified, is that early AI programmes may target easier indications, which would inflate success rates without the technology contributing anything.

What do people inside the field say? The most careful statements come from insiders. Insilico's chief executive has noted that the algorithm identifies a candidate in months while getting it into a patient takes years, costs hundreds of millions, and depends on factors no model can predict. A narrative review's conclusion is that the challenges defining pharmaceutical R&D, target validation, patient heterogeneity, toxicology, pharmacokinetics and trial failure, remain unchanged.

What is the strongest objection to this article? That judging the field by approvals is premature by construction, since drug development takes ten to fifteen years and AI-designed molecules began entering trials around 2020. A second objection is that the molecular-versus-biological split may be tidier than the evidence supports: rentosertib's target was itself model-identified, so if that target proves correct it is direct evidence against the claim that AI has not yet demonstrated anything about biology.

Sources

Primary documents only. Where a claim rests on a single report, the entry says so.

  1. Artificial Intelligence in Drug Discovery (2020-2026): Clinical Successes, Persistent Challenges, and Pharmaceutical Adoption Narrative review for medical professionals The statement that no AI-discovered drug has achieved regulatory approval and that Phase II success rates align with historical industry averages around 40%, alongside the observation that target validation, patient heterogeneity, toxicology and trial failure remain unchanged.
  2. AI Drug Discovery in 2026: 173 Trials, 0 Approvals Science Reader The pipeline count, the Phase I figures of 80 to 90% against 52% traditional and Phase II around 40%, the absence of a standard definition for AI-designed, and the indication-selection concern.
  3. INS018_055 Phase IIa Results IntuitionLabs, on the Nature Medicine publication The rentosertib trial detail: NCT05938920, 71 patients, double-blind placebo-controlled 12 weeks, 98.4 mL FVC improvement at 60 mg with 95% CI 10.9 to 185.9. Note that sources give differing placebo figures for this trial, which the article flags.
  4. AI drug discovery in 2026: between promise and production AskMe The BCG pipeline count and the reference to Jayatunga et al., Drug Discovery Today, June 2024, on how successful AI-discovered drugs are in clinical trials.

Learn the concepts

← All posts