The same PDF says 83% and nobody quotes it
Territory 10 opens on enterprise deployment. The most-quoted statistic in the field is real, measures something much narrower than its use, and is contradicted inside its own source document.
TL;DR. "95% of AI pilots fail" reached Fortune, Forbes, Harvard Business Review, board decks and earnings calls. It comes from The GenAI Divide: State of AI in Business 2025, a self-described preliminary working paper from Project NANDA, based on 150 interviews, 350 employee surveys and 300 public deployments. The figure measures one thing: whether a pilot produced measurable P&L impact within six months, on a sample weighted toward sales and marketing, which the same report identifies as the lowest-ROI area it studied. And the same PDF reports pilot-to-implementation of around 83% for general-purpose chatbots, with more than 80% of organisations having piloted them and nearly 40% reporting deployment. The report's own interview script asks whether respondents saw measurable returns from any AI deployment. Those answers appear nowhere in it. The buried finding that would have been useful: vendor partnerships succeed about 67% of the time against roughly one third for internal builds.
---
Status: established, and the primary document is the source of the criticism. The report is a preliminary working paper. Its stated method, its stated limitations, its own contradicting figures and its own unpublished interview question are all inside it. This article makes no claim that anyone acted in bad faith, and the strongest criticisms of the framing come from people who read the paper carefully rather than from anyone disputing its data.
---
What the number is
*Project NANDA published The GenAI Divide: State of AI in Business 2025 in July 2025, authored by Aditya Challapally, Chris Pease, Ramesh Raskar and colleagues, based on 150 leadership interviews, 350 employee surveys and an analysis of 300 public AI deployments.*
About 5% of pilot programmes achieved rapid revenue acceleration. The rest stalled, delivering little or no measurable impact on profit and loss.
Fortune ran it. Forbes ran it. Harvard Business Review ran it. It appeared on earnings calls and was cited in coverage of market movements. It is, by some distance, the most quoted statistic about enterprise AI.
And it is not fabricated, misattributed or invented. The number is in the document, it means what the document says it means, and the document says what it means quite clearly.
The problem is entirely in the gap between what it measured and what it is used for.
What "fail" meant
Four conditions, all stated in the report.
Deployment beyond pilot. Measurable KPIs. ROI measured six months after the pilot. And elsewhere, a successful tool described as one where users or executives reported marked and sustained productivity or P&L impact.
That is a demanding bar and it is a legitimate one to set. It is also not what "fail" conveys to a reader.
A pilot that improved a process, saved staff time, or produced a working system with no attributable revenue line inside six months counts as a failure by this definition. So would a competent new hire, assessed on the same terms.
The six-month window is the load-bearing choice. The report itself lists it among its limitations, noting it may be too short to judge success at all.
Which means the honest headline is narrower and far less quotable: bespoke, workflow-embedded generative AI struggled to demonstrate profit-and-loss impact within six months, on a sample the authors described as preliminary.
The composition problem
More than half of generative AI budgets in the study went to sales and marketing tools.
The report found the biggest returns in back-office automation: eliminating outsourced business process work, cutting agency costs, streamlining operations.
So the sample is weighted toward the category the study itself identifies as lowest-return.
That is a real finding about resource allocation and it is the one the report leads with in its own analysis. It also means the aggregate failure rate describes where the money went rather than what the technology can do, and those are different claims.
Anyone quoting 95% as a statement about AI capability is quoting a statement about enterprise budgeting.
The figure inside the same document
This is the part that should have ended the discussion.
The same PDF reports that more than 80% of organisations had explored or piloted general-purpose tools such as ChatGPT and Copilot, that nearly 40% reported deployment, and that generic chatbots showed a pilot-to-implementation rate of around 83%.
Eighty-three percent and five percent are in one document, describing two categories, and only one of them travelled.
The report did not find that enterprise AI was failing across the board. It found that custom, embedded, workflow-specific deployments struggled against a hard six-month business-impact test, while general-purpose tools moved from pilot to implementation at a high rate.
Both findings are in the source. One became a headline and the other did not, and the difference is not accuracy. It is that one of them is surprising and the other is not.
The question with no answer printed
The report's appendix contains its interview script. Question 12 asks whether the respondent has seen measurable returns from any AI deployment.
The answers do not appear anywhere in the report.
This was found by Kevin Werbach reading the appendix, and it is the sharpest observation anyone has made about the document. If 150 executives had overwhelmingly answered no, that would be the single most important line in the paper. There is no line.
No inference is available about why, and none is offered here. What can be said is that a question directly bearing on the headline claim was asked and its results were not published, and a reader who has only seen the 95% figure has no way to know that.
And it was hard to obtain
Werbach also documented the retrieval problem. He went to the NANDA site expecting the file and found a form. He filled it in and received nothing. He eventually located the document in someone else's LinkedIn post.
The cover carries the MIT logo. Project NANDA originated at MIT and is not administered by it, which is a distinction that does not survive the phrase "MIT study" and did not survive it in most coverage.
None of this makes the research wrong. It makes the most-cited enterprise AI statistic one that most people quoting it have not read, could not easily have read, and attribute to an institution whose relationship to the work is looser than the branding implies.
That is citation decay with an unusual feature: the source is not obscure or old. It is fourteen months old and was actively difficult to get.
The finding that was actually useful
Buried in the same report: purchasing from specialised vendors and building partnerships succeeded about 67% of the time. Internal builds succeeded about one third as often.
That is an actionable result with a clear mechanism, and it is particularly relevant to financial services and other regulated sectors, where many firms were building proprietary systems.
It is a story about strategy rather than about technology, and it received a fraction of the attention the failure rate did.
Which is a general pattern worth naming. A study contains a striking aggregate and a useful conditional. The aggregate travels because it is quotable; the conditional stays because it requires context to state.
What the prior art says
The broad direction is not new and was not news.
Capgemini found in 2023 that 88% of AI pilots never reached production. S&P Global found 42% of generative AI pilots abandoned.
Both predate the NANDA report. Neither moved a market.
The consistency matters more than the novelty. Three independent measurements, different methods, different years, all finding that most pilots do not reach production. The underlying phenomenon is real and well-supported, which the criticism of the 95% figure does not touch.
What the counter-data says
Q1 2026 figures point the other way on adoption, and they come with their own problems.
72% of enterprises reported at least one AI workload in production, up from 55% in 2024 and 20% in 2020. More than 80% of the Fortune 500 are reported to run agents in production. Average enterprise AI spend reached about $7 million in 2025, projected to rise 65% to $11.6 million in 2026.
And the internal contradiction in that same data is the interesting part. 97% of executives say they are benefiting from AI. 29% report significant organisational ROI.
Those two figures are not in conflict in the way they appear to be. Individual benefit and organisational return are different measurements at different levels, which is the scope problem again: a productivity gain distributed across many people can be real and still not appear as a line in a financial statement.
One database of more than 130 documented enterprise case studies with named results at large firms exists as a counterweight, and its own editor is candid that a database built by looking for successes will find them. That candour is worth more than the database.
The honest position
Kevin Werbach, in a comment under his own critique, stated it better than any of this:
The report's failure to demonstrate widespread failure does not mean deployments are succeeding.
That is the whole thing. A weak study finding failure is not evidence of success. Debunking a statistic returns you to not knowing, not to the opposite conclusion, and a great deal of the commentary in both directions treats it otherwise.
What is actually known: most pilots do not reach production, consistently across several independent measurements over three years. What is not known: whether that reflects the technology, the deployment strategy, the measurement window, or the fact that most pilots of anything do not reach production.
What travels and what stays
This article is the third case in the corpus where a source's own qualifying evidence sat next to the quoted figure and did not move with it.
*The IEA's Energy and AI report contains a worked example showing that all the world's chatbot text queries account for roughly 2% of AI data centre electricity.* The projections from that report circulate constantly. The worked example does not.
Epoch AI's inference price analysis reports declines ranging from 9x to 900x per year depending on the milestone chosen, and notes in the same document that its fastest trends all begin after January 2024 and that benchmark contamination became more common over the same period. The rate circulates. The range and the caveat do not.
And here, 83% sits beside 5% in one PDF.
In none of these cases did anything have to be discovered. The qualifying material was in the source, published by the same authors, in the same document, at the same time. What failed was not research. It was reading.
And the selection is not random. In all three cases the figure that travelled was the more surprising one, the more quotable one, and the one that supported a stronger claim. The material left behind was the material that made the claim conditional, which is precisely the material a reader needs and precisely the material that does not fit in a headline.
Which suggests a cheap and unglamorous check that would have caught all three: before quoting a figure from a study, read what is on the same page. Not the abstract, not the press release, not the coverage. The neighbouring paragraphs, which is where the conditions live.
Why this territory is worth entering
Enterprise deployment is where the claims in the previous four territories get tested against money.
Territory 6 examined nine documented incidents and found no remedy was a better model. Those were failures that reached courts, regulators and parliaments, which is a small and unrepresentative sample by construction.
Territory 8 measured what the buildout costs, and found the composition of the spending largely undisclosed.
Territory 9 measured what cheap generation does to information systems, and found a cost ratio rather than a quality problem.
This territory asks the question those three imply: when an organisation actually deploys one of these systems, with a budget and an owner and a timeline, what happens?
And the answer is currently governed by a statistic almost nobody has read, which is a reasonable place to start.
The measurement problem here is also distinctive. Enterprise deployment sits behind commercial confidentiality rather than behind a technical difficulty. Companies know exactly what their pilots returned. They have the figures, they are not obliged to publish them, and the ones who do publish are self-selecting for having something good to report, which is the survivorship structure this corpus keeps finding.
Which means the honest expectation for this territory is that the good numbers will come from the same places they came from in Territory 8: filings, regulators, and occasional research that somebody funded for reasons of their own.
Three things this establishes
A number can be accurate, well-sourced and still wrong for the use it is put to. The 95% figure is in the document, means what the document says, and is quoted as a claim about capability when it measures six-month P&L attribution on a budget-weighted sample.
The qualifying evidence was in the same file. 83% pilot-to-implementation for general-purpose tools sits alongside the 5% figure. Nothing had to be researched to find it. It had to be read, and the reading is the step that did not happen.
And a preliminary working paper became a settled fact without changing a word. The cover says preliminary. The limitations section lists the sample constraints, the self-selection risk and the short window. All of it survived intact in the document and none of it survived in the citation.
What it does not establish
That enterprise AI is succeeding. The counter-data is adoption data, and adoption is not return.
That the report is bad research. It is a preliminary paper that stated its method, its limits and its contradicting findings, which is more than many widely-cited studies do.
That anyone acted improperly. No claim of that kind is made here, and the sharpest criticisms come from careful readers rather than from anyone disputing the data.
And nothing about any specific company's deployment. Every figure here is aggregate.
What is unresolved
Whether the six-month window is the right one. Nobody has repeated this measurement at eighteen or thirty-six months, which is the obvious follow-up and would settle a great deal.
What Question 12 returned. The answers exist somewhere and are not public.
Whether the vendor-versus-internal gap holds. 67% against roughly 33% is the most actionable finding in the report and has not been independently replicated.
And what the 97%-against-29% gap actually is. Distributed individual benefit that does not aggregate, measurement failure at the organisational level, or both, and no study separates them.
What would settle it
Four specific measurements would resolve most of this, and three are achievable by researchers with no cooperation from anyone.
Repeat the study at eighteen and thirty-six months. The six-month window is the load-bearing choice and the report says so in its own limitations. A pilot judged at six months and again at two years would separate a technology problem from a measurement problem, and nobody has done it.
Replicate the vendor-versus-internal split. 67% against roughly one third is the most actionable claim in the paper, has an obvious mechanism, and rests on a single preliminary sample. It is also the finding an enterprise would actually change behaviour on, which makes replication more valuable than anything else in the document.
Stratify by function. The aggregate is weighted toward sales and marketing, which the study identifies as lowest-return. Separate rates for back-office automation, customer operations, engineering and sales would turn one misleading number into four useful ones, and the underlying data to do it already exists.
And publish the Question 12 answers. That requires the authors and nobody else.
The fourth is the interesting one, because it is the only item on the list that cannot be obtained by doing more work. Everything else is a study somebody could run. That answer exists, was collected, and is simply not public, which is the same shape as every disclosure gap in the previous territory: the party who could resolve the question has no obligation to.
The counter-argument
Criticising the 95% figure has become its own genre, and that genre has an interest. Vendors, consultancies and practitioners all benefit from the claim that enterprise AI is working better than reported, and several of the most prominent critiques come from parties selling into the market. The critique is correct and the incentive is real, which is the standard this corpus applies elsewhere.
The report's defenders have a point about the bar. Demanding measurable P&L impact within six months is exactly what boards demand, so a study measuring against that standard is measuring something real rather than something artificially strict.
The 83% counter-figure may be weaker than presented. Pilot-to-implementation for a general-purpose chatbot is a much lower bar than for an embedded workflow system, since deploying ChatGPT to staff requires little integration. Comparing the two rates directly, as this article does, is not comparing like with like.
And the retrieval complaint may be overstated. A working paper being hard to download is ordinary for preliminary research, and the document was available, if awkwardly. Treating distribution friction as a substantive problem risks importing a standard that most working papers would fail.
The pattern this territory will test
Four territories have each found that the intuitive explanation was the wrong one, and enterprise deployment is where that claim faces its most commercially motivated audience.
The intuitive account of enterprise AI failure is that the models are not good enough yet. It is comfortable, it implies a solution arriving on its own, and it locates the problem outside the organisation.
The NANDA report's own most useful finding points elsewhere. Vendor partnerships at 67% against internal builds at roughly one third is not a statement about model capability. The same models were available to both groups. What differed was who did the integration work, who owned the workflow knowledge, and who had done it before.
That is the shape Territory 7 found in robotics, where difficulty did not predict feasibility and task shape did. It is the shape Territory 6 found in the incident record, where no remedy was a better model. And it is the shape Territory 9 found, where the damage was a cost ratio rather than output quality.
If the pattern holds here, the pilots that failed did not fail on capability. They failed on workflow fit, on integration ownership, on measurement windows chosen before anyone knew what to measure, and on budgets allocated to the function with the lowest return.
All four of those are organisational variables, and all four are things an organisation controls.
The falsification test is straightforward and worth stating in advance. If a subsequent study stratifies by function, controls for integration approach and measurement window, and still finds capability to be the dominant explanatory variable, the framing this corpus has applied across four territories is wrong and should be discarded rather than defended. No such study currently exists in either direction.
The short version
"95% of AI pilots fail" comes from a self-described preliminary working paper, based on 150 interviews, 350 employee surveys and 300 public deployments, and it measures one thing: whether a pilot produced measurable P&L impact within six months, on a sample weighted toward sales and marketing, which the same report identifies as its lowest-return category.
The same PDF reports around 83% pilot-to-implementation for general-purpose chatbots, with more than 80% of organisations having piloted them. Both figures are in one document. One travelled and one did not, and the difference is that one is surprising.
Its own interview script asks whether respondents saw measurable returns from any deployment. Those answers are not printed. The document was hard to obtain, and the MIT logo on the cover belongs to a project that originated there and is not administered by it.
The finding that would have been useful was buried: vendor partnerships succeeding about 67% of the time against roughly one third for internal builds, which is a statement about strategy rather than technology.
The direction is not new. Capgemini found 88% of pilots never reaching production in 2023; S&P Global found 42% abandoned. Neither moved a market, and their consistency with the NANDA finding is the strongest evidence that the underlying phenomenon is real.
And the honest position is the narrow one. A weak demonstration of failure is not a demonstration of success. Debunking a statistic returns you to not knowing.
Common questions
Where does the 95% figure come from? From The GenAI Divide: State of AI in Business 2025, a preliminary working paper published in July 2025 by Project NANDA, based on 150 leadership interviews, 350 employee surveys and an analysis of 300 public AI deployments. It reports that about 5% of pilot programmes achieved rapid revenue acceleration while the rest stalled with little or no measurable P&L impact.
Is the number wrong? No. It is in the document, it means what the document says, and the document states its method clearly. The problem is the gap between what it measured and what it is quoted for. Success was defined as deployment beyond pilot, with measurable KPIs, and ROI assessed six months after the pilot, described elsewhere as marked and sustained productivity or P&L impact. That is a demanding and legitimate bar, and it is not what "fail" conveys to a general reader.
What does the same report say that contradicts the headline? That more than 80% of organisations had explored or piloted general-purpose tools such as ChatGPT and Copilot, that nearly 40% reported deployment, and that generic chatbots showed a pilot-to-implementation rate of around 83%. Both the 5% and the 83% figures are in the same PDF, describing different categories. The report found that custom, workflow-embedded deployments struggled against a six-month business-impact test, not that enterprise AI was failing generally.
What is the issue with Question 12? The report's appendix contains its interview script, and Question 12 asks whether the respondent has seen measurable returns from any AI deployment. The answers appear nowhere in the report. If 150 executives had overwhelmingly answered no, that would be the most important line in the paper. No inference about why is available and none is offered, but a reader who has seen only the headline figure has no way to know the question was asked.
Why does the sample composition matter? Because more than half of generative AI budgets in the study went to sales and marketing tools, and the report itself identifies back-office automation as the highest-return area. So the aggregate failure rate is weighted toward the category the study says returns least. That is a genuine finding about enterprise resource allocation, and it means the number describes where money went rather than what the technology can do.
What was the useful finding? That purchasing from specialised vendors and building partnerships succeeded about 67% of the time, while internal builds succeeded about one third as often. That is actionable, has a clear mechanism, is particularly relevant to regulated sectors building proprietary systems, and received a small fraction of the attention the failure rate did. Aggregates travel because they are quotable; conditionals stay behind because they need context to state.
Does other research agree that most pilots fail? Broadly yes, and this is the part the criticism does not touch. Capgemini found in 2023 that 88% of AI pilots never reached production, and S&P Global found 42% of generative AI pilots abandoned. Both predate the NANDA report and neither moved a market. Three independent measurements using different methods across three years all find most pilots not reaching production, which makes the underlying phenomenon well supported even where the famous number is misused.
So is enterprise AI working or not? Unknown, and that is the honest answer rather than an evasion. Adoption is clearly rising: 72% of enterprises report at least one AI workload in production, up from 55% in 2024, with spend projected to rise 65% in 2026. Return is much less clear: 97% of executives say they are benefiting while 29% report significant organisational ROI, which are different measurements at different levels rather than a contradiction. As Kevin Werbach put it, the report's failure to demonstrate widespread failure does not mean deployments are succeeding. Debunking a statistic returns you to not knowing.
Sources
Primary documents only. Where a claim rests on a single report, the entry says so.
- The GenAI Divide: State of AI in Business 2025 Project NANDA, July 2025, preliminary working paper The primary document: 150 leadership interviews, 350 employee surveys, 300 public deployments, the 5% figure and its four success conditions, the 83% pilot-to-implementation rate for general-purpose tools, the 67% against one third vendor split, and the interview script in the appendix.
- Myth Number 2: MIT Showed That 95% of AI Pilots Fail NewMR The close reading of the success definition, the general-purpose tool figures inside the same report, and the point that the study found one type of bespoke deployment struggling rather than enterprise AI failing.
- The MIT 95% AI Failure Rate Is Not What You Think Medium Kevin Werbach's findings on Question 12 and the retrieval problem, the MIT NANDA administration distinction, the Capgemini and S&P Global prior figures, and the observation that failure to demonstrate failure is not evidence of success.
Further reading
The primary literature behind the claims above, drawn from the concept entries this post links to, so a claim carries the same source here as it does there.
- Raji et al. (2021), AI and the Everything in the Whole Wide World Benchmark — how a specific measurement becomes a general claim through restatement. :: https://arxiv.org/abs/2111.15366 Citation Decay
- Lipton & Steinhardt (2018), Troubling Trends in Machine Learning Scholarship, arXiv:1807.03341 — the mechanisms by which claims outrun their evidence in a literature. :: https://arxiv.org/abs/1807.03341 Citation Decay
- Wong et al. (2021), External Validation of a Widely Implemented Proprietary Sepsis Prediction Model — the difference between overall performance and performance on the population that matters. :: https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2781307 Scope Boundary
- Pineau et al., reproducibility programme work, and Bouthillier et al. (2021), Accounting for Variance in Machine Learning Benchmarks — why a literature of reported results describes what was submitted. :: https://arxiv.org/abs/2103.03098 Survivorship Bias
Related articles
- Zero-click is 60%, or 22.4%, from one providerTerritory 9 opens on what cheap generation does to information. Every method agrees the traffic is falling. None agrees on how far, and the headline metric differs threefold within one dataset.
- The EU delayed the part with no standardsThree AI Act obligations take effect today and the widely reported headline says the opposite. What was deferred, what was not, and why the split falls exactly where it does.
- Every chatbot query on earth is 2% of AI's powerTerritory 8 opens on the numbers behind AI's physical footprint, and on the arithmetic in the IEA's own report that almost nobody quotes.
- Three years in: no disruption, and one 20% holeEvery aggregate measure of AI's labour effect shows continuity. One within-firm comparison shows a fifth of a cohort gone. Both are well evidenced.