Home/Blog/Evaluation & evidence/The pilot ran on data the production system will never see

The pilot ran on data the production system will never see

Enterprise AI pilots are built on a curated slice, pre-cleaned, with limited users and manual review. Then production data arrives. That sequence explains the failure rate without invoking model capability at all.

TL;DR. A Cloudera and Harvard Business Review Analytic Services survey of 1,574 enterprise IT leaders, published March 2026, found 7% saying their data is completely ready for AI adoption. A separate Cloudera index of 1,270 IT leaders found 96% integrating AI into core business processes, 85% claiming a clear data strategy, and around 80% admitting their initiatives are constrained by limited data access. Those describe the same organisations. Readiness figures range from 5% to 19% depending on definition. Gartner projects 60% of AI projects lacking AI-ready data will be abandoned through 2026. And the mechanism is specific: a pilot selects a manageable slice, pre-cleans it, limits the user base and manually reviews outputs before each review. It succeeds under borrowed conditions and breaks when production data arrives, which means the pilot never tested the thing that determines production.

---

Status: consistent direction, definitional spread, and almost entirely vendor-commissioned. The largest surveys here are run by data platform companies whose product is the remedy. The Cloudera and HBR instrument states its sample, fielding window and weighting, which puts it above most of this literature and does not remove the interest. Figures are attributed throughout.

---

The readiness figures and their spread

Cloudera with Harvard Business Review Analytic Services surveyed 1,574 enterprise IT leaders and published in March 2026 that only 7% say their data is completely ready for AI adoption.

Snowflake's enterprise research found only 7% of organisations with the majority of their unstructured data in an AI-ready state.

Dun & Bradstreet reported 97% of organisations running active AI initiatives and 5% believing their data can support AI at enterprise scale.

AI Markets Group benchmarked 2,048 enterprise decision-makers and found 87% using AI, 70% having adopted generative AI, and 19% fully data-ready.

Five percent, seven percent, seven percent, nineteen percent. The spread is definitional: completely ready, majority of unstructured data ready, can support AI at enterprise scale, and fully data-ready are four different bars, and each survey chose its own.

The direction is unanimous across four independent instruments and the magnitude is not established. A reader quoting any single figure is quoting a definition somebody else picked.

The paradox in the same instrument

Cloudera's Data Readiness Index surveyed 1,270 IT leaders at companies with more than 1,000 employees, fielded 22 January to 3 March 2026, weighted to the GDP of surveyed countries.

96% report integrating AI into core business processes.

85% say they have a clear data strategy.

And around 80% admit their AI and data initiatives remain constrained by limited data access across environments.

Those three figures are from one survey of one population. Cloudera calls it an AI readiness illusion, where adoption outpaces the foundation required to deliver impact.

The 85% against 80% pair is the interesting one. Most organisations believe they have a data strategy and most report being blocked by data access. Both can be true if a strategy exists as a document rather than as an implemented estate, which is what the readiness figures independently suggest.

The mechanism

This is the part the pilot literature does not state and it explains the failure rate without reference to models.

A typical enterprise AI pilot is built on borrowed conditions. A team selects a manageable slice of data, pre-cleans it, limits the user base, and manually reviews model outputs before each stakeholder review.

The model performs well. The pilot passes its success criteria.

Then production arrives: the full estate rather than the slice, uncleaned, with the whole user base and no manual review in the loop.

Which means a successful pilot is evidence about a curated environment and is used as evidence about an uncurated one. That is external validation in the form this corpus documented in the sepsis case, where a model validated by its developer performed differently when checked independently on the population that mattered.

And it reframes the 95% figure entirely. A pilot that showed no measurable return may have been a pilot whose conditions never existed outside it. The technology did not fail a test. The test did not resemble the deployment.

What actually blocks it

A survey of 650 enterprise technology leaders published in March 2026 attributes 89% of the failures preventing pilots from reaching production to five causes.

Integration complexity, 63%. Output quality degradation at scale, 58%. Insufficient monitoring infrastructure, 54%. Unclear organisational ownership, 49%. Domain-specific training data gaps, 41%.

Four of the five are organisational or infrastructural. One concerns data content. None is about model capability.

A separate consultancy audit of 47 enterprise engagements from 2022 to 2025 states it more bluntly: the model was almost never the blocker. The named blockers were inconsistent metric definitions, no certified source for the entities the model needed to reason about, and pipelines never built to feed an inference workload.

And the reported time split supports it. Teams commonly find 60 to 70% of AI project time going to data preparation, which is a familiar figure from the analytics era and is now attached to a different set of expectations.

Why agents made old debt visible

Something changed in 2026 and it is worth naming precisely.

Data debt is not new. Fragmented entity resolution, undocumented pipelines and metric definitions that differ per tool have been ordinary enterprise conditions for decades. Analytics tolerated them because a human sat between the data and the decision.

An analyst who receives a report with two conflicting customer counts notices, asks, and reconciles. The reconciliation is invisible, unbilled and constant, and it is what made the estate workable.

An agent does not do that. It acts on what it retrieves.

Which is why one definition of AI-ready data is specifically data an agent can act on without supervision, as distinct from a data platform, which delivers a place to put data.

So the requirement did not get harder. The tolerance disappeared. The same estate that supported dashboards for ten years fails an agent, and the failure looks like an AI problem because the agent is the new element.

That is the binding constraint moving without anyone deciding it should. The constraint was always the data layer and a human was absorbing it.

What the ready minority did differently

The consistent finding across sources is unglamorous and pre-technical.

McKinsey's 2025 State of AI found high performers nearly three times as likely to have fundamentally redesigned workflows, with strong technology and data infrastructure among the practices most associated with meaningful value.

The consultancy audit names five data-layer failures: untrusted metric definitions, fragmented entity resolution, pipelines that batch where streaming is needed, governance documented but not enforced, and a metric layer that re-derives the same measure differently in each tool. Its claim is that fixing any three ships most pilots.

And the obstacles reported by leaders are the same ones: siloed data at 56% and absence of a clear data strategy at 44%, with 73% saying their organisation should prioritise AI data quality more than it currently does.

None of that is an AI programme. It is data engineering, entity resolution and governance, which are decades-old disciplines with no novelty value and no budget line labelled AI.

Which explains the incentive problem. An organisation can fund a pilot or fund a catalogue. The pilot demonstrates in eight weeks and the catalogue demonstrates nothing, and the catalogue is what determines whether the pilot survives contact with production.

What the pilot removed, item by item

Setting the four conditions against their production equivalents makes the gap concrete rather than rhetorical.

Pilot conditionProduction conditionWhat the pilot could not test
Curated sliceFull estateEntity resolution across systems
Pre-cleaned dataWhatever existsQuality handling at volume
Limited user baseEveryoneQuery variety and edge cases
Manual output reviewNoneWhether errors are caught

The right-hand column is the finding. Each removed condition corresponds to a specific untested capability, and all four are the capabilities that determine whether a deployment works.

Which means a pilot's pass is close to uninformative about production unless the pilot deliberately retained at least some of the production conditions, and the standard pilot design retains none.

The fix is not complicated and it is unpopular. Run the pilot on a slice of uncleaned production data, with no manual review of outputs, and with a user group that did not design it. It will fail more often, earlier, and for less money, which is the entire value and also the reason it is rarely done.

And it explains a pattern the pilot literature reports without explaining. Deployments that fail rarely fail at launch. They fail some weeks later, when the manual review that nobody documented as part of the system stops happening, which is when the pilot's fourth removed condition finally takes effect.

Where this leaves the corpus's own claim

Article 159 stated a falsification test in advance, and it is worth checking against this evidence rather than waiting for a more convenient moment.

The test was this: if a study stratifies by function, controls for integration approach and measurement window, and still finds capability the dominant explanatory variable, the framing this corpus has used across five territories is wrong and should be discarded.

No such study exists. What exists here is a five-cause survey of 650 leaders where four of five causes are organisational, and a consultancy audit of 47 engagements stating the model was almost never the blocker. Neither is the controlled study the test asked for, and both point away from capability.

So the framing survives and it has not been tested. A survey of self-reported failure causes is not a controlled comparison, and organisations have an obvious reason to attribute failure to infrastructure rather than to having chosen the wrong problem.

The honest position is the one this corpus reached about the deferred-standards question in the AI Act article: the evidence is consistent with the framing and does not confirm it, and a framework surviving because the discriminating study does not exist has been left alone rather than validated.

What would settle it is specific. Two comparable deployments, same organisation, same function, one on a readied data estate and one not, with the model held constant. That is an A/B test a single large enterprise could run, and its absence is the same absence this territory has found in every article.

Three things this establishes

A successful pilot may be evidence about nothing that will persist. Curated slice, pre-cleaned data, limited users, manual review. Every one of those is removed at production, and the pilot measured the system with all of them present.

Adoption and readiness are different measurements and both are being reported as AI progress. 96% integrating AI into core processes and 7% with data ready describe one population. The first figure travels because it sounds like momentum.

And the constraint was always the data layer. What changed is that analytics tolerated an unreliable estate because a human reconciled it invisibly, and an agent acts on what it retrieves. The requirement did not rise. The absorption stopped.

What it does not establish

That data readiness is the only blocker. The five-cause survey puts integration complexity first and monitoring third, neither of which is data content.

That the readiness figures are precise. Four instruments give 5%, 7%, 7% and 19% on four different definitions, and none is wrong.

That fixing data guarantees success. The consultancy claim that fixing three of five failures ships most pilots is a vendor's account of its own engagements, and no controlled comparison exists.

And nothing about model capability. This article says capability was not the binding constraint in the cases measured, which is a claim about what limited these deployments rather than about what models can do.

What is unresolved

Whether readiness investment pays. The organisations pouring budget into catalogues and lineage cannot demonstrate return either, and the argument that they are positioned for a later cycle is a prediction.

What the definitional bar should be. Until one is agreed, the 5% to 19% spread persists and every figure remains a choice.

Whether the agent tolerance argument holds empirically. It is mechanically plausible, widely repeated, and no study measures how much reconciliation human analysts were actually performing.

And whether pilots could be designed to test production conditions. Running a pilot on uncurated data with no manual review would be a genuine test and would fail far more often, which is a reason it does not happen.

Why the definitional spread is the useful part

Five percent, seven percent, seven percent, nineteen percent. The instinct is to pick one or average them. Both are wrong and the spread itself carries information.

Each bar describes a different operational reality.

Nineteen percent, fully data-ready, is the loosest and probably describes organisations with a working platform, some governance and reasonable quality. Enough to run a pilot.

Seven percent, completely ready for AI adoption, is stricter and closer to the state where a deployment survives contact with production.

Seven percent with a majority of unstructured data ready is a specific technical claim about indexing and retrievability, which matters because more than 80% of enterprise data is unstructured and most readiness programmes address the structured part.

And five percent, believing data can support AI at enterprise scale, is the strictest and includes a judgement about scale rather than about state.

Read together they describe a gradient rather than a contradiction. Roughly a fifth of organisations can run something. Roughly one in fourteen can deploy it. Roughly one in twenty believe it scales.

That gradient is more useful than any single figure and no source presents it, because each publisher has one number and an interest in it being memorable.

It also predicts where the failures land. An organisation in the 19% but not the 7% will pass its pilot and fail its deployment, which is the exact population the failure statistics are counting, and the exact sequence this article describes.

Which is the general lesson about definitional spread, and it applies well beyond this subject. A range across four instruments with four stated bars is not noise to be averaged away. It is a rough measurement of how the population thins as the requirement tightens, and it is available for free to anyone who declines to pick a favourite number.

The counter-argument

Almost every figure here comes from a company selling data infrastructure. Cloudera, Snowflake and the consultancies all sell the remedy for the problem they measured, and a finding that enterprises need better data foundations is the most commercially convenient conclusion available to them. The Cloudera and HBR instrument states its method, which is more than most, and it does not make the interest disappear.

Curated pilots are not obviously wrong. Testing a new system on a controlled slice before exposing it to a full estate is ordinary engineering practice, and the alternative, running an unproven system on production data with no review, is worse. The criticism is really that pilot results are over-interpreted, which is a reporting failure rather than a design failure.

The data-readiness framing can excuse anything. Any failed deployment can be attributed to insufficient data maturity after the fact, and the claim is difficult to falsify because no organisation ever has perfect data. A theory that explains every failure explains none of them, and this literature has that shape.

And the agent argument may be overstated. Retrieval systems can be designed to flag conflicts, request clarification and refuse to act on ambiguous data, and several do. Treating an agent as inherently unable to reconcile is a statement about current implementations rather than about the technology.

The short version

Cloudera with Harvard Business Review Analytic Services surveyed 1,574 IT leaders and found 7% saying their data is completely ready for AI. Snowflake found 7% with a majority of unstructured data ready, Dun & Bradstreet found 5% believing their data can support AI at enterprise scale, and AI Markets Group found 19% fully data-ready across 2,048 decision-makers. Four instruments, four definitions, one direction.

And a separate Cloudera index of 1,270 leaders found 96% integrating AI into core processes, 85% claiming a clear data strategy, and around 80% blocked by limited data access. One survey, one population, three figures.

The mechanism is that a pilot runs on borrowed conditions: a curated slice, pre-cleaned, with limited users and manual review before each stakeholder session. All four are removed at production, so a pilot that passed measured a system that will not exist.

Which reframes the failure rate. A survey of 650 leaders attributes 89% of pilot failures to integration complexity at 63%, quality degradation at scale at 58%, insufficient monitoring at 54%, unclear ownership at 49% and training data gaps at 41%. Four of five are organisational. A consultancy audit of 47 engagements puts it plainly: the model was almost never the blocker.

And agents did not create the debt, they stopped absorbing it. An analyst receiving two conflicting customer counts reconciles them invisibly and constantly. An agent acts on what it retrieves, which is why an estate that supported dashboards for a decade now fails, and why the failure looks like an AI problem.

Common questions

How many enterprises have AI-ready data? Between 5% and 19% depending on the definition used. Cloudera with Harvard Business Review Analytic Services surveyed 1,574 IT leaders and found 7% saying their data is completely ready for AI adoption. Snowflake found 7% with the majority of unstructured data in an AI-ready state. Dun & Bradstreet found 5% believing their data can support AI at enterprise scale. AI Markets Group, across 2,048 decision-makers, found 19% fully data-ready. The direction is unanimous across four independent instruments and the magnitude is not established.

What is the readiness illusion? Cloudera's term for adoption outpacing foundations. Its index of 1,270 IT leaders at companies over 1,000 employees, fielded from 22 January to 3 March 2026, found 96% integrating AI into core business processes and 85% claiming a clear data strategy, while around 80% admitted their initiatives remain constrained by limited data access. Those figures come from one survey of one population, and the strategy-against-access pair is consistent with a strategy existing as a document rather than as an implemented estate.

Why do pilots succeed and production deployments fail? Because they run under different conditions. A typical pilot selects a manageable slice of data, pre-cleans it, limits the user base and manually reviews outputs before each stakeholder review. The model performs and the pilot passes. Production removes all four: the full estate, uncleaned, with the whole user base and no human in the loop. A successful pilot is therefore evidence about a curated environment being used as evidence about an uncurated one.

What actually blocks deployments? A March 2026 survey of 650 enterprise technology leaders attributes 89% of failures to five causes: integration complexity at 63%, output quality degradation at scale at 58%, insufficient monitoring infrastructure at 54%, unclear organisational ownership at 49%, and domain-specific training data gaps at 41%. Four of the five are organisational or infrastructural. A consultancy audit of 47 engagements from 2022 to 2025 states that the model was almost never the blocker, naming inconsistent metric definitions, fragmented entity resolution and pipelines never built for inference workloads.

Why did this become visible in 2026 specifically? Because the tolerance disappeared rather than the requirement rising. Fragmented data, undocumented pipelines and inconsistent metric definitions have been ordinary enterprise conditions for decades, and analytics tolerated them because a human sat between the data and the decision. An analyst who receives two conflicting customer counts notices and reconciles, invisibly and constantly. An agent acts on what it retrieves. The same estate that supported dashboards for ten years fails an agent, and the failure looks like an AI problem because the agent is the new element.

What do the organisations that succeed do differently? Unglamorous, pre-technical work. McKinsey's 2025 State of AI found high performers nearly three times as likely to have fundamentally redesigned workflows, with strong data infrastructure among the practices most associated with value. The named data-layer fixes are trusted metric definitions, entity resolution producing a single customer or patient record, pipelines that stream where inference requires it, enforced rather than documented governance, and one metric layer rather than per-tool re-derivation. None of that is an AI programme, which is the incentive problem: a pilot demonstrates in eight weeks and a data catalogue demonstrates nothing.

How reliable are these figures? Directionally consistent and commercially interested. Cloudera, Snowflake and the consultancies quoted all sell the remedy for the problem they measured, which is the most convenient possible finding for them. The Cloudera and HBR instrument states its sample, fielding window and weighting, which puts it above most of this literature without removing the interest. No independent survey of enterprise data readiness exists.

What is the strongest objection to this framing? That data readiness explains too much. Any failed deployment can be attributed to insufficient data maturity after the fact, no organisation ever has perfect data, and a theory compatible with every outcome is not doing explanatory work. A second objection is that curated pilots are ordinary engineering practice and the real error is over-interpreting their results, which is a reporting failure rather than a design failure.

Sources

Primary documents only. Where a claim rests on a single report, the entry says so.

  1. The Data Readiness Index: Understanding the Foundations for Successful AI Cloudera, April 2026 The 1,270 IT leader survey with stated fielding window and GDP weighting: 96% integrating AI into core processes, 85% claiming a data strategy, around 80% constrained by data access. Commissioned by a data platform company.
  2. AI-Ready Data: Why Enterprise AI Pilots Fail in Production Enqurious The borrowed-conditions mechanism, the five-cause survey of 650 enterprise technology leaders from March 2026, and the Snowflake finding of 7% with a majority of unstructured data AI-ready.
  3. Only 7% of Enterprises Have AI-Ready Data Luiz Neto The reconciliation of the 5%, 7% and 19% figures across Dun & Bradstreet, Cloudera with HBR Analytic Services and AI Markets Group, and the argument that agents made pre-existing data debt visible.
  4. The 2026 Enterprise Guide to AI-Ready Data NX1 The Cloudera and Harvard Business Review Analytic Services survey of 1,574 IT leaders finding 7% completely ready, the Gartner projection on abandonment, and the 60 to 70% of project time spent on data preparation.

Further reading

The primary literature behind the claims above, drawn from the concept entries this post links to, so a claim carries the same source here as it does there.

  • Wong et al. (2021), External Validation of a Widely Implemented Proprietary Sepsis Prediction Model — the canonical case, and the source of the 7% figure. :: https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2781307 External Validation
  • Raji et al. (2021), AI and the Everything in the Whole Wide World Benchmark — why a developer-selected evaluation cannot license a general claim. :: https://arxiv.org/abs/2111.15366 External Validation
  • International Energy Agency (2025), Energy and AI — the infrastructure lead times against which compute demand is set. :: https://www.iea.org/reports/energy-and-ai Binding Constraint
  • International Federation of Robotics, World Robotics 2025 — an installed base showing where physical automation was and was not constrained. :: https://ifr.org/worldrobotics/report-2025 Binding Constraint

Learn the concepts

← All posts