AI in science: 380,000 predicted, 736 actually made
AlphaFold won a Nobel Prize and did not reduce the rate of experimental structure determination. A materials model predicted 380,000 stable compounds and 736 have been synthesised.
TL;DR. The single most successful AI-for-science result in history won the 2024 Nobel Prize in Chemistry and released structures for over 200 million proteins, a roughly 1,500-fold increase over everything laboratories had characterised in decades. An economic study of its effect found the rate of experimental structure determination almost unchanged. Researchers used the predictions to complement experiments rather than replace them, and shifted their attention toward proteins that had no structural information at all. A materials model predicted 2.2 million candidate structures and flagged around 380,000 as stable; outside laboratories have physically synthesised 736. Prediction is not discovery, and the gap between the two is where almost all of the confusion in this area lives.
---
In July 2021 structural biologists gained access to hundreds of thousands of AI-predicted protein structures, effectively overnight. The database has since grown past 200 million structures and is used by more than two million researchers across 190 countries. In 2024 the work received the Nobel Prize in Chemistry, the first awarded for an AI-enabled scientific breakthrough.
Economists then studied what happened to the field.
The rate of experimental structure determination remained almost unchanged.
That is the most informative sentence in the literature on AI and science, and it is not the sentence anyone expected. A tool that predicts, at near-experimental accuracy, the thing a field spent decades determining experimentally, did not reduce the amount of experimental determination.
What changed was what the experiments were about. Researchers used predicted structures to facilitate and complement experimental work rather than substitute for it, and basic research increased most on proteins that previously had no structural information at all. The tool did not replace the labour. It redirected it toward questions that had been out of reach.
The number that should be quoted more than it is
If AlphaFold is the success case, materials discovery is the case that shows what a headline number conceals.
A materials model predicted 2.2 million candidate crystal structures and identified roughly 380,000 as stable, including around 52,000 novel lithium-ion conductors. This was widely reported as the discovery of hundreds of thousands of new materials.
As of the paper reporting it, external laboratories had physically synthesised 736.
That is 0.19% of the stable set.
The number is not a scandal and it is not evidence the work was worthless. It is a description of what a computational screen produces: a vastly wider funnel of things worth trying, not a set of finished materials. Each of the 380,000 is a hypothesis about thermodynamic stability. A hypothesis is a claim about what an experiment would show, and until someone runs the experiment it remains one.
The failure is in the reporting verb. "Discovered 380,000 materials" and "predicted 380,000 candidates worth synthesising" describe the same result and imply completely different things about what exists in the world.
Prediction, discovery and validation are three different events
The distinction that resolves most disagreement in this area, and it applies well beyond materials.
Prediction is a model producing a claim. It is cheap, it scales, and its output volume tells you about the model's throughput rather than about the world.
Validation is an experiment testing that claim. It is expensive, it does not scale, and it is where the physical world is consulted. The ratio between prediction volume and validation volume is the single most useful statistic about any AI-for-science claim, and it is almost never reported. In the materials case it is roughly 500 to 1.
Discovery is a validated claim that turned out to be both true and useful. Most validated predictions are true and unremarkable. A stable compound nobody needs is a stable compound.
The rhetorical move that causes trouble is using the word "discovery" for the first of these. Once you separate them, most disputes about whether AI is transforming science become disputes about which of the three someone is counting, and those are resolvable.
Where the results are real
Three areas have produced results that survive this scrutiny, and it is worth being specific about why.
Structure prediction. The strongest case in the field. The successor model extended predictions from single protein chains to complexes involving nucleic acids and small molecules, with reported improvements of at least 50% over prior methods. Around 40% of new structures deposited in the main structural database in 2024 and 2025 involved AI-assisted techniques alongside experimental ones. This is a real capability with real adoption.
Design with experimental closure. Methods for designing proteins that do not exist in nature, where the design is then synthesised and tested in a laboratory. The distinguishing feature is that the loop closes: a claim is made and then physically checked. That is discovery in the strict sense, and it is a much smaller body of work than the prediction literature.
Screening that narrows an intractable space. The materials case belongs here when described honestly. A search space of billions reduced to 380,000 worth attempting is a genuine contribution to a field where the binding constraint is deciding what to try next.
Notice what these share: the model produces candidates and the physical world adjudicates. The successful pattern in science is the same as the successful pattern in medicine and law, which is that the system proposes and something outside it disposes.
What has not happened
Worth stating plainly, because the absence is informative.
No disease has been cured. Several AI-originated drug candidates have entered human clinical trials, which is a real milestone and a new one. Entering trials is not the same as working; most candidates entering trials do not become drugs, and that base rate applies here as everywhere.
No autonomous discovery. Models propose experiments and hypotheses in narrow, data-rich domains. Humans and instruments verify every result. The 2026 position is an accelerated discovery loop with people in it, not a scientist that runs on its own.
And no reduction in experimental burden, at least in the field with the strongest tool. That is the AlphaFold finding, and it should temper expectations everywhere else: the most successful case did not save the experiments, it changed which experiments were worth doing.
What it did to the scientists
The second-order effects are better documented here than in any other domain, because economists find field-wide shocks irresistible, and they are more interesting than the capability results.
Researchers moved away from what the model predicts well. One study found experimental structural biologists pivoting away from proteins that the model handles reliably, particularly where downstream demand is limited. That is rational: there is no reward for experimentally confirming what everyone can already predict. It also means the tool reshapes the research agenda rather than only the research method.
Adoption concentrated among the already-productive. Highly productive structural biologists were more likely to adopt, and the study reports this exacerbating citation polarisation between more and less cited researchers. A tool that is free to everyone still accrues disproportionately to those positioned to use it, which is the Matthew effect operating through a technology that looks like a leveller.
Laboratories changed who they hired. Labs led by life-science specialists responded by hiring more computer scientists. The composition of a field shifted because of a tool, which is a slower and more consequential change than the tool itself.
None of these effects is captured by any benchmark, and all of them are larger in aggregate than the accuracy improvement that produced them.
The replication question nobody has answered
Science has spent fifteen years confronting a replication crisis. AI arrived in the middle of it, and the two problems interact in a way that has not been worked through.
Machine learning results are harder to replicate than the results they are used to produce. A biology experiment has a protocol. A model has a protocol, a training set, a random seed, a hardware configuration, a library version and a set of hyperparameters that may not have been recorded. Fixing the seed does not deliver bitwise reproducibility, because parallel floating-point arithmetic sums identical values in different orders.
So a scientific claim resting on a model inherits two replication problems rather than one.
And the incentives point the wrong way. A paper reporting that a model found something is publishable. A paper reporting that someone tried to reproduce the model and could not is much less so, and it requires compute the reproducing team may not have.
The specific worry is a class of paper that is easy to produce and hard to check. Train a model on a dataset, report that it identified a pattern, publish. If the dataset was contaminated, if the split leaked, if the metric was chosen after seeing the results, none of that is visible in the paper, and the field consuming the result frequently lacks the machine learning expertise to ask.
Three practices would address most of it, none of them novel, all of them borrowed from fields that solved this earlier. Pre-register the analysis so the metric is fixed before the result is seen. Publish the code and the data split, not just the description. And report the number of configurations tried, because a result selected from a hundred attempts is a different claim from a result obtained on the first.
The reason this belongs in an article about scientific discovery rather than an article about methodology: a discipline that cannot check its computational results will accumulate them anyway, and the correction will arrive later and cost more than the practice would have.
How to read an AI-for-science claim
Six questions, and the first two do most of the work.
Was anything physically made or measured? If the result is entirely computational, it is a hypothesis set. That can be extremely valuable and it is not a finding about the world.
What is the ratio of predictions to validations? 2.2 million to 736 is a legitimate result described honestly and a misleading one described as discovery. Ask for both numbers.
Was the validation done by someone else? External synthesis is a much stronger signal than in-house confirmation, for the same reason external validation matters in medicine.
What was the baseline method? "Faster than the previous computational approach" and "faster than the experiment" are very different claims, and the second is the one people hear.
Did the field's behaviour change? Adoption by working scientists is a harder test than benchmark performance and a better one. Two million users is evidence; a leaderboard position is not.
And what happened to the experimental rate? If a predictive tool is truly substituting for experiments, the experiments should decline. In the best-documented case they did not, which suggests complementarity is the normal outcome and substitution is the exception.
What is unresolved
Whether complementarity generalises. The finding that experimental rates held steady comes from one field with one exceptionally good tool. Whether that reflects something general about how prediction interacts with experiment, or something specific to structural biology's incentives, is unknown and matters enormously for forecasting.
Whether the 736 becomes 7,360. Synthesis is slow and the predictions are recent. The honest position is that the validation rate is a running total rather than a final one, and the question is what fraction of a large candidate set ever gets attempted at all.
What happens to the tacit knowledge. Experimental structure determination carries craft knowledge that is transmitted by doing. If the field redirects toward problems where prediction suffices, some of that capacity may not be replaced, and it is the capacity needed to check the predictions.
And whether the concentration effects compound. If adoption accrues to the already-productive and shifts citations toward them, a tool distributed free to everyone could still widen the gap between institutions. Whether that is transitional or self-reinforcing has not been established.
The counter-argument
Complementarity is not disappointment. The AlphaFold result is often read as deflationary and it is not. A tool that leaves the experimental rate unchanged while redirecting it toward previously unreachable problems has increased the total amount of science being done, not failed to reduce the cost of the existing amount. Measuring it by labour saved is measuring the wrong quantity.
736 is a timing artefact more than a hit rate. Synthesis takes years, funding follows interest slowly, and the predictions are recent. Comparing a five-year computational output against a two-year experimental response and calling the ratio a finding is unfair to both.
The Nobel is a fact and it is not a small one. A committee not known for enthusiasm about computation judged this the most significant advance in its field. Any framework for evaluating AI in science that produces a sceptical read on the one result the discipline itself has recognised at that level should be examined for whether the framework is wrong.
And the counterfactual is unknowable. The experimental rate held steady, and nobody knows what it would have done otherwise. Structural biology could have been contracting for unrelated reasons, in which case a flat rate is a substantial positive effect being read as no effect at all.
The short version
AlphaFold released structures for over 200 million proteins, roughly a 1,500-fold increase over decades of laboratory characterisation, is used by more than two million researchers in 190 countries, and won the 2024 Nobel Prize in Chemistry, the first for an AI-enabled scientific breakthrough. An economic study of its effect on structural biology found the rate of experimental structure determination almost unchanged. Researchers used predictions to complement experiments, and increased basic research most on proteins that previously had no structural information.
The tool did not replace the labour. It redirected it toward questions that had been unreachable.
The materials case shows what a headline conceals. A model predicted 2.2 million candidate crystal structures and flagged around 380,000 as stable, reported widely as the discovery of hundreds of thousands of materials. External laboratories have physically synthesised 736, about 0.19%. That is a legitimate result described honestly and a misleading one described as discovery, and the difference is entirely in the verb.
Three events get conflated. Prediction is a model producing a claim, which is cheap and scales. Validation is an experiment testing it, which is expensive and does not. Discovery is a validated claim that turned out to matter. The ratio of the first to the second is the most useful statistic about any AI-for-science claim and is almost never reported; in the materials case it is roughly 500 to 1.
The second-order effects are larger than the capability results and no benchmark captures them. Researchers pivoted away from proteins the model predicts well, since confirming what everyone can already compute earns nothing. Adoption concentrated among the already-productive, exacerbating citation polarisation. Laboratories hired computer scientists.
And what has not happened is as informative as what has. No disease cured, though candidates have entered trials. No autonomous discovery, since humans and instruments verify every result. And no reduction in experimental burden in the field with the best tool available, which suggests complementarity is the normal outcome and substitution is the exception.
Common questions
Has AI actually made any scientific discoveries? It has produced results that led to validated discoveries, and the distinction matters. Protein structure prediction won the 2024 Nobel Prize in Chemistry and is used by over two million researchers. Designed proteins that do not occur in nature have been synthesised and tested successfully. But most reported AI discoveries are predictions awaiting validation: a materials model flagged around 380,000 stable candidates and external laboratories have physically synthesised 736.
What is the difference between predicting and discovering? Prediction is a model producing a claim, which is cheap and scales with compute. Validation is an experiment testing that claim, which is expensive and does not scale. Discovery is a validated claim that turned out to be both true and useful. Reporting frequently uses "discovery" for the first, which is why headline numbers about hundreds of thousands of new materials describe a hypothesis set rather than things that exist.
Did AlphaFold replace laboratory work? No, and this is the most surprising finding in the area. An economic study of its release found the rate of experimental structure determination almost unchanged. Researchers used predicted structures to facilitate and complement experimental determination rather than substitute for it, and increased basic research most on proteins that had no prior structural information. The tool redirected the labour rather than reducing it.
How many of the AI-predicted materials have actually been made? As of the paper reporting the result, external laboratories had physically synthesised 736 of roughly 380,000 predicted stable structures, from an initial 2.2 million candidates. That is about 0.19%. Synthesis is slow and the predictions are recent, so the figure is a running total rather than a final hit rate, but it is the number that should accompany any claim about hundreds of thousands of discovered materials.
Has AI cured any diseases? No. Several AI-originated drug candidates have entered human clinical trials, which is a genuine milestone and new. Entering trials is not the same as working, and the base rate for candidates entering trials and becoming approved drugs is low and applies here as elsewhere. No AI-originated therapy has completed the path to demonstrated clinical benefit.
Is AI doing science autonomously? No. Models propose experiments and hypotheses in narrow, data-rich domains, and humans and instruments verify every result. What exists in 2026 is an accelerated discovery loop with people in it rather than a system that conducts research independently. The successful pattern is the same as in medicine and law: the model proposes and something outside it adjudicates.
How did AI change scientists' behaviour? More than it changed their output, and the effects are well documented. Experimental structural biologists pivoted away from proteins the model predicts reliably, since confirming a computable result earns little. Adoption concentrated among already-productive researchers, which one study found exacerbating citation polarisation. And laboratories led by life-science specialists responded by hiring more computer scientists, changing the composition of the field.
How should I evaluate a claim that AI discovered something? Ask whether anything was physically made or measured, since a purely computational result is a hypothesis set. Ask for the ratio of predictions to validations, since 2.2 million to 736 is honest when stated and misleading when summarised. Ask whether validation was external. Ask what the baseline was, since faster than a previous computation and faster than an experiment are different claims. And ask whether the experimental rate in that field changed, since in the best-documented case it did not.
Related articles
- How to check an AI claim before you believe itA vendor advertised a hallucination rate below 0.001% with no evidence behind it, and a state attorney general investigated. Six questions separate a claim you can act on from a number someone typed.
- AI in education: students feel twice the gain they getA 2026 review of 66 randomised trials found AI tutoring improved satisfaction by 0.93 and confidence by 0.91, against 0.53 for knowledge. The authors rated all three very low certainty.
- Testing a system that answers differently every timeForty percent of organisations hit significant quality regressions within ninety days of deploying an LLM application. Exact-match assertions are the cause: they reject valid answers and occasionally accept wrong ones.
- AI in government: 126 use cases, 65 not made publicOne agency reported 126 active AI use cases and auditors found the inventory still incomplete, with tools contracted to build criminal cases missing from it entirely.