The pilot failed and the staff deployed it anyway
Enterprise AI is measured by what organisations sanctioned. A separate literature measures what their employees actually use, and the two describe the same companies without ever being set against each other.
TL;DR. IBM's Cost of a Data Breach found 20% of organisations suffered a breach linked to shadow AI, adding as much as $670,000 to the average breach cost, with 97% of affected organisations having no AI access controls and 63% having no AI governance policy at all. Only 17% have technical controls preventing uploads of confidential data to public AI tools. Gartner's survey of 302 security leaders found 69% suspect or have evidence of prohibited use. And the finding nobody states is that this literature and the enterprise pilot literature describe the same organisations. One says most AI initiatives fail to show measurable return. The other says employees adopted AI so effectively that the security function cannot see it. Both are true, and the pilot statistic measures what was sanctioned rather than what is used.
---
Status: one strong anchor, wide surrounding variance. IBM's annual breach study has a stated methodology and a long series. The surrounding figures come overwhelmingly from vendors selling shadow AI detection, and they disagree with each other by wide margins, which is stated where used rather than averaged away.
---
The anchor figures
*IBM's Cost of a Data Breach reported that one in five studied organisations experienced a breach linked to shadow AI*, defined as unsanctioned AI tools adopted without IT or security oversight.
Those incidents added as much as $670,000 to the average breach cost, and disproportionately exposed customer personally identifiable information and intellectual property.
97% of the organisations involved had no AI access controls in place.
63% of organisations had no AI governance policy at all.
And only 17% have technical controls preventing employees uploading confidential data to public AI tools. The other 83% rely on training, warning emails, or nothing.
Separately, Gartner surveyed 302 cybersecurity leaders between March and May 2025 and found 69% suspecting or holding evidence that employees use prohibited public generative AI tools.
Those are the numbers this article rests on. IBM's study has a published method and a multi-year series; Gartner's has a stated sample and window. Everything else in this subject is looser and is treated as such below.
Where the surrounding numbers disagree
The variance is worth showing rather than resolving, because the spread is the honest picture.
Employee adoption of unsanctioned tools is reported at 57%, at 78%, and at 98% depending on source, with the highest figure describing organisations having any exposure rather than individuals using tools, which is a different measurement wearing the same sentence.
Employees entering confidential data into public tools appears at 27%, at 33%, and at 36%.
Netskope reported 47% of generative AI users accessing tools through personal accounts, bypassing enterprise controls entirely.
And Cyberhaven Labs reported mid-level employees out-using their managers by roughly 3.5 times.
Almost every one of these is published by a company selling detection or governance tooling. The direction is consistent across independent vendors and the magnitudes should be read as indicative. A range from 57% to 98% for what is nominally one quantity is a definitional problem rather than a measurement dispute, and none of the publishers states which definition they used.
The finding these two literatures produce together
This is the part that neither literature states, because they are written by different people for different buyers.
The enterprise pilot literature reports that most AI initiatives fail to show measurable return within six months. It measures sanctioned deployments: projects with owners, budgets, timelines and success criteria.
The shadow AI literature reports that employees adopted AI so thoroughly and so fast that most organisations cannot see it, and a fifth have suffered a breach because of it.
Both describe the same companies in the same period.
Which produces a reading available from neither alone. In a substantial number of organisations, the official AI programme did not demonstrate return and the unofficial one demonstrated enough value that people risked their jobs to keep using it.
One analysis puts the mechanism plainly: the exposure is driven by capable people doing legitimate work, and the organisation is not fighting carelessness but productivity.
And the survey evidence supports that reading directly. Reported reasons for unsanctioned use centre on unapproved tools offering better functionality than approved alternatives, and on the absence of a capable approved option. Mid-level employees out-using managers by 3.5 times is not a compliance failure profile. It is an adoption profile.
What the pilot statistic was actually measuring
The 95% figure measured whether a pilot produced measurable profit-and-loss impact within six months, on a sample weighted toward sales and marketing.
It did not measure whether anyone in the organisation was getting value from AI.
An employee saving forty minutes a day drafting with a tool nobody approved contributes nothing to that measurement. There is no project, no owner, no line item, and no attribution. The productivity exists and the accounting does not.
Which is a scope boundary of an unusually consequential kind, because both figures are then used to answer the same question. "Is enterprise AI working?" gets answered with a number about governed projects, in a period where most of the actual usage was ungoverned.
Nothing here says the pilot studies were wrong. They measured what they said they measured. The error is downstream, where a figure about sanctioned initiatives became a figure about the technology.
Why banning does not work, and what the record says instead
The shadow IT precedent is directly on point and the industry has run this experiment before.
In the 2010s employees routed around IT with consumer cloud tools. Security responded with access brokers and data loss prevention, discovering unsanctioned applications and blocking risky uploads. The lesson recorded from that cycle was that bans alone failed, and visibility plus sanctioned alternatives worked.
Generative AI compressed the same cycle into roughly twenty-four months.
Blocking pushes usage to personal devices and personal accounts, where the data still leaves and the organisation no longer sees it. The 47% figure for personal-account access is what that looks like when it has already happened.
The measures the record supports are unremarkable. Provide a capable approved tool, because 27% of unsanctioned users cite better functionality as the reason. Establish visibility before control. And write a policy, since 63% of organisations have none, which makes every subsequent conversation about enforcement premature.
The cost lands where it usually does
Cost externality applies here in an unusually clean form.
The employee pasting a contract into an unapproved assistant captures the benefit: forty minutes saved, work delivered on time, a manager satisfied.
The organisation carries the exposure: the breach cost, the regulatory position, the intellectual property in a third party's system.
And the security function carries the burden of detection, without having chosen the tools, set the policy, or been consulted about whether an approved alternative exists.
None of the three parties is behaving unreasonably. The employee is doing the job faster, the organisation is buying productivity, and the security team is responding to a decision made elsewhere. The cost simply does not sit with whoever created it, which is the condition under which a burden grows until somebody exits.
The distinctive feature here is that the exposure is created by success. A tool nobody used would create no shadow AI problem. The measured harm is a function of how well the unofficial deployment worked.
The two pictures, set against each other
Nobody publishes this table, which is the point of drawing it.
| Enterprise pilot literature | Shadow AI literature | |
|---|---|---|
| What it counts | Sanctioned projects | Unsanctioned usage |
| Who commissions it | Consultancies, research bodies | Security vendors |
| Headline | Most initiatives show no return | Most organisations cannot see the use |
| Implied conclusion | Enterprise AI is not working | Enterprise AI adoption is out of control |
| Population | Enterprises discussing AI programmes | Enterprises running detection tooling |
Two industries, two buyers, two conclusions, one set of companies.
The consultancy buyer wants to know whether to invest more. The security buyer wants to know whether to buy controls. Neither has a commercial reason to commission the study that would reconcile them, which would ask how much value an organisation is getting from AI in total, governed and ungoverned, and no such study exists.
The bottom row is where the joint reading is weakest and it belongs in the table rather than in a footnote. Enterprises willing to discuss their AI programme and enterprises running shadow AI detection are overlapping populations rather than identical ones. This article's central claim depends on that overlap being substantial, which is plausible and unestablished.
What the table does establish regardless of the overlap is that the two most-quoted framings of enterprise AI in 2026 are produced by industries with opposite commercial interests, and that a reader encountering either in isolation is getting the half that somebody paid to have measured.
What an organisation could measure
Three things, from data most already hold, none requiring a vendor.
Total AI usage against sanctioned AI usage. Network telemetry and expense reports both reveal it. The ratio is the single number that would tell an organisation whether its pilot results describe its AI position or a fraction of it.
Reasons, gathered without consequences attached. The survey evidence says people use unapproved tools because approved ones are worse or absent. That is checkable inside one organisation in a week, and it distinguishes a procurement gap from a discipline problem, which determines the entire response.
And the value estimate nobody makes. An organisation that knows its pilots returned nothing and does not know what its ungoverned usage returned has measured the smaller half and concluded about the whole.
The obstacle is not technical. Asking employees what unapproved tools they use, in an environment where the answer is punishable, produces no data. Asking in an amnesty produces the inventory and the reasons together, and costs nothing but the decision to not punish the answers.
Which is the same conclusion this corpus reaches in every territory, arriving from a new direction. The party with the data is the organisation itself. There is no obligation to look, and the incentive to look is weaker than the discomfort of what would be found.
Three things this establishes
Measuring sanctioned deployment measures governance, not adoption. A pilot study counts projects. An organisation where every pilot failed and every employee uses AI daily scores as a failure, and describing that as the technology not working is a category error.
The variance in this literature is definitional and unresolved. 57%, 78% and 98% are presented as the same quantity by different vendors with different denominators. A figure with a threefold spread and no stated definitions is a description of the market for detection tools, and the IBM and Gartner figures are the ones with methods attached.
And the exposure grows with the value. Every measure of shadow AI harm is downstream of employees finding the tools useful enough to use without permission. A governance response that treats this as a discipline problem is addressing the symptom of a procurement gap.
What it does not establish
That shadow AI is harmless. A fifth of organisations suffering a linked breach, at $670,000 of additional cost, is a serious finding from a credible source.
That employees are right to bypass controls. The intellectual property and personal data exposures are real, and the fact that a behaviour is understandable does not make it safe.
That the pilot studies should have measured usage. They measured what they set out to measure, and a study of governed initiatives is a legitimate thing to conduct.
And nothing about any specific organisation's exposure. That depends on an inventory this article cannot perform.
Where the shadow IT comparison breaks
The precedent is invoked constantly, including above, and the disanalogy is rarely stated. It should be.
Shadow IT meant data sitting in an unsanctioned application. A spreadsheet in a personal cloud account was in the wrong place, accessible to the wrong parties, and outside backup and retention policy. Serious, and recoverable in principle: the file could be located, deleted, and the account closed.
Data placed in a generative AI tool may not be recoverable in that sense.
Three differences matter. It may transit or train an external model, depending on the provider's terms and the account type, which is a different disposition from storage. It may be irreversible: a file can be deleted, and a model that has been trained on something cannot straightforwardly be untrained. And the boundary is invisible to the user, since the interface for a consumer account and an enterprise account with data protection terms is frequently identical.
IBM's own framing acknowledges the continuity and the escalation: shadow AI mirrors the rise of shadow IT a decade ago, with far higher stakes.
Which affects the remedy in one specific way. The shadow IT playbook was discover, then block or sanction, then remediate. The remediation step is weaker here, because discovering that a contract was pasted into a consumer account eighteen months ago does not undo it.
That strengthens rather than weakens the article's practical conclusion. If remediation is limited, prevention matters more, and the only prevention with a record is providing a tool people prefer. Blocking produces personal-account usage, which is precisely the disposition where the data is least recoverable.
And it sharpens the timing. Shadow IT tolerated a slow response because the damage accumulated in place. Here the exposure is completed at the moment of use, which is why the 63% of organisations with no policy is a more urgent figure than it looks.
What is unresolved
What fraction of AI value in enterprises is currently ungoverned. Nobody has measured it, and both literatures would have to be redesigned to find out.
Whether approved alternatives actually reduce shadow use. It is the recommendation everywhere and the before-and-after evidence is thin.
What the definitions are. The 57% to 98% spread cannot be narrowed without publishers stating denominators, and none does.
And whether the shadow IT lesson transfers. Cloud applications sat outside the organisation and did not train on the data placed in them. The comparison is used constantly and the disanalogy is rarely stated.
The corpus has been running on commissioned evidence
This article has spent several sections identifying whose commercial interest shaped which measurement, so the same test applied here is overdue.
Territory 10 has now used, across five articles, one preliminary working paper from a project with a commercial position, vendor pricing comparisons in which each publisher won, security surveys from companies selling remediation, benchmark analyses from firms selling evaluation, and shadow AI figures from detection vendors.
The exceptions are the ones worth naming. A model developer's system card publishing its own attack success rates. A peer-reviewed benchmark paper introducing a metric that made its authors' field look worse. IBM's breach study, which has a method and a series and predates the product category it now describes. And the AI Act itself, which is a legal instrument rather than a finding.
That is four reliable anchors across five articles, and everything else is context assembled from parties with positions.
Which is not a confession of failure. In most of these subjects the alternative was not a disinterested measurement; it was no measurement, and this corpus has said so about other people's evidence often enough to accept it about its own.
But it does bound what Territory 10 can conclude. A territory built substantially on commissioned evidence can describe what the commissioning parties measured, can identify where their questions do not meet, and cannot say what is true in the gap.
The specific gap here is the largest. Nobody funds a study of total organisational AI value, governed and ungoverned together. Both halves are measured by industries with reasons to measure only their half, and the reconciling question sits in the space between two well-funded literatures, unasked.
Naming that is the honest limit of this article and, on the evidence assembled so far, of the territory.
The counter-argument
Reading shadow AI as evidence of successful adoption is a considerable stretch. People also paste data into tools that waste their time, and unsanctioned use measures availability and habit at least as much as it measures value. A tool being used is weak evidence that it works, which is a standard this corpus applies rigorously elsewhere and relaxes here.
Nearly this entire subject is vendor literature. Strip out the companies selling shadow AI detection and what remains is IBM's breach study and one Gartner survey. The article's own framing depends on adoption figures it simultaneously describes as unreliable, which is an uncomfortable position.
The productivity framing may excuse a genuine failure. Pasting client contracts and source code into third-party systems is a serious control failure regardless of the motive, and characterising it as capable people doing legitimate work risks making a security problem sound like a procurement oversight.
And the two literatures may not be comparable at all. The pilot studies sample enterprises willing to discuss AI programmes; the shadow AI figures sample organisations running detection tooling. Setting them side by side assumes a shared population that neither establishes, and the joint reading this article builds on rests on that assumption.
The short version
IBM found 20% of organisations suffered a breach linked to shadow AI, adding as much as $670,000 to average breach cost, with 97% of those affected having no AI access controls and 63% having no AI governance policy at all. Only 17% have technical controls on uploads to public tools. Gartner found 69% of 302 security leaders suspecting or evidencing prohibited use.
The surrounding literature disagrees with itself. Employee adoption of unsanctioned tools appears at 57%, 78% and 98%, confidential data entry at 27%, 33% and 36%, and almost every figure comes from a company selling detection.
And the finding neither literature states is that this and the enterprise pilot literature describe the same organisations. One reports that most AI initiatives showed no measurable return in six months. The other reports that employees adopted AI so effectively the security function cannot see it.
So in a substantial number of companies, the official programme failed to demonstrate value and the unofficial one demonstrated enough that people used it without permission. The pilot statistic measured governed projects with owners and budgets. An employee saving forty minutes a day with an unapproved tool appears nowhere in it.
The precedent is shadow IT and the recorded lesson is that bans failed while visibility plus sanctioned alternatives worked. Blocking moves usage to personal accounts, which 47% of generative AI users already use.
And the cost lands where it usually does. The employee captures the benefit, the organisation carries the exposure, and the security function carries the detection burden without having chosen any of it. The exposure is created by the tools being good enough to use.
Common questions
What is shadow AI? The use of AI tools without IT or security review, approval or visibility. It covers employees pasting data into consumer chatbots, developers using AI coding assistants connected to production systems, AI browser plugins, personal-account access to work tasks, and unauthorised agents operating inside corporate environments.
What are the reliable numbers? IBM's Cost of a Data Breach reported that 20% of studied organisations experienced a breach linked to shadow AI, adding as much as $670,000 to the average breach cost and disproportionately exposing customer personal data and intellectual property. Of those affected, 97% had no AI access controls, and 63% of organisations had no AI governance policy at all. Only 17% have technical controls preventing confidential uploads to public AI tools. Gartner separately surveyed 302 cybersecurity leaders between March and May 2025 and found 69% suspecting or evidencing prohibited public generative AI use.
Why do the other figures vary so much? Because they measure different things under the same words and mostly come from companies selling detection tooling. Employee adoption of unsanctioned tools appears at 57%, 78% and 98%, where the highest figure describes organisations having any exposure rather than individuals using tools. Confidential data entry appears at 27%, 33% and 36%. A threefold spread on a nominally single quantity with no stated denominators is a definitional problem rather than a disagreement about the world.
What is the connection to enterprise pilot failure rates? They describe the same organisations and nobody sets them side by side. The pilot literature reports that most AI initiatives fail to show measurable profit-and-loss impact within six months, measuring sanctioned projects with owners, budgets and success criteria. The shadow AI literature reports that employees adopted AI thoroughly enough that most organisations cannot see it. In a substantial number of companies, the official programme showed no return while the unofficial one delivered enough value that people used it without permission.
Does that mean the pilot studies were wrong? No. They measured what they said they measured, and studying governed initiatives is legitimate. The error is downstream, where a figure about sanctioned projects becomes a figure about the technology. An employee saving forty minutes a day with an unapproved tool contributes nothing to a profit-and-loss attribution, because there is no project, no owner and no line item. The productivity exists and the accounting does not.
Does blocking AI tools work? The recorded precedent says no. In the 2010s employees routed around IT with consumer cloud tools, and the lesson from that cycle was that bans alone failed while visibility plus sanctioned alternatives worked. Blocking pushes usage onto personal devices and personal accounts, where the data still leaves and the organisation loses visibility. Netskope reported 47% of generative AI users already accessing tools through personal accounts.
What does the evidence support instead? Providing a capable approved tool, since a common stated reason for unsanctioned use is that unapproved tools offer better functionality than approved alternatives. Establishing visibility before attempting control. And writing a policy at all, given that 63% of organisations have none, which makes conversations about enforcement premature.
What is the strongest objection to this article's framing? That reading unsanctioned use as evidence of value is a stretch. People use tools that waste their time, and adoption measures availability and habit at least as much as it measures usefulness, which is a standard this corpus applies strictly elsewhere. A second objection is that the two literatures may not sample the same population: pilot studies survey enterprises willing to discuss AI programmes, while shadow AI figures come from organisations running detection tooling, and the joint reading assumes a shared population that neither establishes.
Sources
Primary documents only. Where a claim rests on a single report, the entry says so.
- Cost of a Data Breach Report IBM The anchor figures: 20% of studied organisations with a breach linked to shadow AI, as much as $670,000 added to average cost, disproportionate exposure of customer personal data and intellectual property, and the framing that shadow AI mirrors shadow IT with higher stakes.
- What Is Shadow AI: Risks, Detection and Prevention NeuralTrust The Gartner survey of 302 cybersecurity leaders finding 69% suspecting or evidencing prohibited use, the 97% without access controls, and the Netskope figure of 47% accessing tools through personal accounts. Published by a governance vendor.
- Shadow AI Risks: The 2026 Security Guide Verax The 17% of organisations with technical controls on confidential uploads, the Cyberhaven finding that mid-level employees out-use managers by roughly 3.5 times, and the framing that the exposure is driven by capable people doing legitimate work. Published by a security vendor.
- Shadow AI: 20% of Breaches, $670K Cost Shattered The shadow IT precedent and the recorded lesson that bans alone failed while visibility plus sanctioned alternatives worked, and the spread of employee adoption figures across circulating surveys.
Further reading
The primary literature behind the claims above, drawn from the concept entries this post links to, so a claim carries the same source here as it does there.
- Raji et al. (2021), AI and the Everything in the Whole Wide World Benchmark — operationalisation determining what a measurement can support. :: https://arxiv.org/abs/2111.15366 Scope Boundary
- Wong et al. (2021), External Validation of a Widely Implemented Proprietary Sepsis Prediction Model — the difference between overall performance and performance on the population that matters. :: https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2781307 Scope Boundary
- Stenberg (2026), The end of the curl bug-bounty — evaluation burden externalised onto an unpaid maintainer until the programme closed. :: https://daniel.haxx.se/blog/2026/01/26/the-end-of-the-curl-bug-bounty/ Cost Externality
Related articles
- Refactoring fell from 25% to 3.8%Survey evidence says AI improves code quality. Repository telemetry says the opposite. They are measuring different things, and the gap between them is where the productivity went.
- Fraud doubles in 18 months. Retraction takes 40.The scientific literature is the one place in this territory with a real record, and the record shows the correction machinery growing at less than half the rate of the thing it corrects.
- Eight subjects, one ratio, and it is not qualityTerritory 9 closes. Across eight subjects the damage came from a cost ratio inverting rather than from bad output, and in every case detection failed while friction worked.
- The 12% was non-inferior, and P was 0.41MASAI is the best-evidenced AI deployment in medicine and the headline everyone quoted describes a result the trial did not claim.