Home/Blog/Foundations/Nine subjects, and the model was almost never the blocker

Nine subjects, and the model was almost never the blocker

Territory 10 closes. Across nine enterprise deployment subjects the binding constraint was organisational in every one, the evidence was commissioned in almost all of them, and the reconciling study exists nowhere.

TL;DR. Nine subjects, and one finding underneath all of them: the constraint was organisational, not technical, in every case where anyone measured. A survey of 650 leaders put four of its five top failure causes outside model capability. A consultancy audit of 47 engagements stated the model was almost never the blocker. The most-quoted failure statistic measured six-month profit attribution on a sample weighted toward its own lowest-return function. Meanwhile 20% of organisations had a breach from tools nobody approved, 7% say their data is ready, 81% expect multiple model providers, and only 17% have technical controls on what leaves the building. And the territory's evidence base is almost entirely commissioned: four reliable anchors across nine subjects, with everything else published by parties selling the remedy they measured.

---

Status: synthesis. No new factual claims. Every figure appears in one of the nine Territory 10 articles with its own sourcing and caveats, linked where used. Source quality varies more here than in any previous territory and is stated per row below.

---

The nine

SubjectThe measured thingWhat actually bound itEvidence quality
Pilots95% show no six-month P&LDefinition and sample compositionPreliminary paper, self-described
Pricing$0.50 to $2.00 per outcomeWho defines the billable eventVendor comparisons, own docs
Reliability61% at one attempt, 25% at eightWhich metric gets reportedPeer-reviewed benchmark
Security4.7% at one attempt, 63% at 100Capability the agent is grantedDeveloper system cards
RegulationThree duties live, high-risk deferredWhether standards existedLegal instrument
Shadow AI20% breached, 63% no policyAbsence of an approved alternativeIBM breach study
Code quality59% say better, refactoring at 3.8%Which instrument you askSurvey against telemetry
Data7% ready, 96% deployingThe estate underneathVendor-commissioned surveys
Lock-in19 to 34% switching costScaffolding, not the APIGateway vendors

Finding one: the constraint was organisational every time

Nine subjects, and in the ones where anyone looked for a cause, capability was not it.

A March 2026 survey of 650 enterprise technology leaders attributed 89% of pilot failures to five causes: integration complexity, quality degradation at scale, insufficient monitoring, unclear ownership, and domain data gaps. Four of five are organisational or infrastructural.

A consultancy audit of 47 engagements from 2022 to 2025 states it flatly: the model was almost never the blocker.

And the pattern repeats where nobody framed it as a cause. Agent reliability turns on scaffolding, which is an engineering choice. Prompt injection exposure turns on what capability an agent is granted, which is a permissions decision. Lock-in turns on prompts and evaluations, which are artefacts a team wrote. Shadow AI turns on the absence of an approved tool.

Which is the fourth consecutive territory to reach this. Territory 6 found no documented incident fixed by a better model. Territory 7 found task shape predicting feasibility where difficulty did not. Territory 9 found a cost ratio rather than an output-quality problem.

Four territories, one shape, and it is now a standing claim rather than an observation. The intuitive explanation, that the technology is not good enough yet, has not been the operative variable in any subject this corpus has examined with numbers attached.

Finding two: almost all of this evidence was commissioned

This is the territory's most uncomfortable result and it constrains everything above.

Four reliable anchors across nine subjects. A peer-reviewed benchmark paper introducing a metric that made its own field look worse. Developer system cards publishing attack success rates against interest. IBM's breach study, which has a method and a series predating the product category. And the AI Act, which is a legal instrument rather than a finding.

Everything else was published by a party selling the remedy. Pricing comparisons by competitors who each win their own table. Switching costs by gateway vendors. Data readiness by data platform companies. Security exposure by remediation sellers. And the most-quoted statistic in the entire subject comes from a self-described preliminary working paper.

That is commissioned framing at territory scale. The evidence base has the shape of its buyers rather than of the questions that matter.

The consequence is specific and worth stating rather than softening. This territory can describe what commissioning parties measured. It can identify where their questions fail to meet. It cannot say what is true in the gap, and the gap is where most of the interesting questions live.

Finding three: the reconciling study never exists

In five of the nine subjects, two literatures describe the same organisations and never meet.

Sanctioned deployment against actual usage. Pilot studies count governed projects; shadow AI counts what employees do. Both describe the same companies, and nobody funds a study of total organisational AI value, governed and ungoverned together.

Perception against telemetry. Surveys report improved code quality; repository analysis reports refactoring collapsing. The study that asks a developer how a change felt and then tracks the churn on that exact commit would settle it, requires one organisation and one quarter, and does not exist.

Benchmark against production. τ-bench reports pass^k; no enterprise publishes its own. Every organisation running agents has the logs.

Attempts against exposure. System cards scale injection success to a hundred attempts; nobody publishes how many attempts a production agent receives.

And readiness against outcome. Two comparable deployments, same organisation, same function, one on a readied data estate and one not. An A/B test a single large enterprise could run.

Five missing studies, all cheap, all requiring data the organisation already holds. None has a commercial sponsor, because in each case the party who could run it is the party who would look worse.

Finding four: the self-report is always more favourable

Three instances, all in the same direction, and the mechanism is not dishonesty.

Developers estimated they were 20% faster and were measured 19% slower on the same tasks in the same study.

97% of executives report benefiting from AI while 29% report significant organisational return.

59% of developers report improved code quality while refactoring falls to 3.8% and churn rises.

And a fourth datapoint from a workforce survey sharpens it. SHRM's 2026 report found 26% of individual contributors saying their trust in AI declines with heavy reliance, against 17% of managers and 9% of directors and above.

Trust falls with proximity to the work. The people furthest from the output are the most confident about it, which is the same gradient the self-report finding describes, measured across an organisation rather than across time.

The mechanism is partial visibility: effort saved is felt immediately, and cost that arrives later, lands on someone else or spreads across a system is felt by nobody. A director reports the benefit their organisation captured. A contributor reports the correction they performed.

Finding five: the cost lands on whoever did not choose it

Every subject in this territory has a party absorbing a burden they did not create.

The reviewer absorbs the check on generated code, with senior engineers reporting 20 to 35% more review time.

The security function absorbs shadow AI detection without having set the policy or been consulted on whether an approved tool exists.

The maintainer absorbs churn, and the person inheriting a duplicated block in eighteen months absorbs the rest.

The advertiser pays above clean rates for inventory every quality metric called premium.

And the buyer absorbs a switching cost created by engineers solving real problems well.

In none of these did the absorbing party decline, because in most of them declining means declining the job.

Which is cost externality as the territory's second unifying mechanism, alongside the organisational constraint. The two are connected: a burden nobody chose is also a burden nobody measures, which is why the constraint stays invisible until somebody exits.

Finding six: the same measurement failure, in five forms

The nine subjects share a measurement structure, and setting the forms side by side shows it is one failure wearing five costumes.

A figure measuring one thing is used to answer another. The pilot statistic measures six-month profit attribution on governed projects and answers "does enterprise AI work". Different question, same number.

A figure at one attempt is used to describe many. Reliability at pass@1 and injection success at a single attempt are both reported, and both describe a situation nobody is in.

A figure from the operator is used where the artefact was needed. Developers, executives and directors all report; repositories, telemetry and churn all disagree.

A figure with a self-chosen definition is used as though standard. Resolution defined by the party invoicing it. Readiness defined four ways producing 5%, 7%, 7% and 19%. Adoption defined three ways producing 57%, 78% and 98%.

And a figure from a party selling the remedy is used as though disinterested, which describes most of the numbers in this territory.

Five forms, and the common property is that each figure is accurate. None of these is fabrication, misattribution or error. Every one is a correct measurement of something adjacent to the question being asked, which is why none of them can be caught by checking the arithmetic or verifying the citation.

Which is the practical lesson the territory produces, and it is a question rather than a rule: what exactly was measured, at what unit of repetition, by whom, under whose definition, and paid for by which party? Five questions, none requiring expertise, all answerable from the source document, and their absence explains most of what this territory found.

What this corpus would need to be wrong

Four territories have now produced the same shape, which is either a finding or a habit, and the difference matters enough to state the discriminating evidence.

The claim is that the operative variable has been structural rather than technical in every subject examined with numbers attached: procedure in the incident record, task shape in robotics, cost ratio in generation, and organisation in deployment.

Three things would break it.

A subject where capability was the binding constraint and the structural account fails. Not a subject where capability mattered, which is common, but one where the organisational, procedural and definitional explanations were tested and did not hold. None has appeared, and this corpus has not gone looking for one, which is a meaningful admission.

A controlled study contradicting the survey pattern. The five missing studies would each provide one, and the data-readiness A/B is the cleanest: two comparable deployments, model held constant, one on a readied estate. If the readied one fails equally, the data-constraint account is wrong.

Or a demonstration that the selection is doing the work. Nine subjects chosen for having numbers, in a corpus written to check numbers, will over-represent subjects where measurement is the interesting problem. The test is whether a subject selected by someone else, for other reasons, produces the same shape.

None of these has been run and the framework has not been confirmed. As the AI Act article put it about a different question, a framework surviving because the discriminating evidence does not exist has been left alone rather than validated, and recording that at the close of a territory is more useful than the territory's conclusions.

What the territory does not show

That enterprise AI is failing. Adoption is near-universal, DORA finds real throughput gains, and one neocloud's growth figures describe customers paying for compute. The failure statistic measures a specific and demanding bar.

That capability is irrelevant. The claim is that capability was not the binding constraint in the cases measured, which says what limited these deployments rather than what models can do.

That the commissioned evidence is wrong. It is frequently the only measurement of anything, and the alternative in most of these subjects was silence.

And that nine subjects are representative. They were selected for having numbers, which selects for domains where somebody had a reason to measure.

What would settle it

SubjectThe missing measurementWho could run it
PilotsSame study at 18 and 36 monthsAny researcher
ReliabilityPass^k on a production agent by task typeAny enterprise
SecurityAttempts per agent per dayAny enterprise
Code qualityPaired survey and telemetry, same commitsOne engineering org
DataA/B on readied against unreadied estateOne large enterprise
Shadow AITotal usage against sanctioned usageAny organisation
Lock-inOne untuned portability testAny team

Six of seven require no cooperation from anyone outside the organisation.

None requires a vendor, a new tool or a research budget. Most require somebody to run a query and write down the answer.

Which makes the absence a statement about incentive rather than difficulty, and the incentive is consistent: in each case the party best placed to measure is the party who would be embarrassed by the result.

What this territory corrected about itself

A synthesis that grades other people's evidence should list its own corrections, and Territory 10 produced four worth recording.

The byline was wrong. Every article on this site carried a personal byline, which implied one person had written 164 articles averaging over three thousand words in under three months. That was not what happened. The first two articles keep the personal byline because they were written that way; the rest now state that they were drafted with AI assistance from primary sources retrieved during the work, reviewed before publication, with one person accountable. The About page previously said "working solo" and "one person read the filings", and both had stopped being true.

Two well-supported concepts were refused. Domain distance had four appearances and was declined as scope boundary applied to domains. Borrowed conditions fitted cleanly and was declined because the external validation node's own one-liner already said it. Refusing a candidate that fits, because an existing node covers it, is the harder call, and the alternative is a graph with a node for every domain an existing mechanism operates in.

One concept was added on thin evidence and labelled thin. Undecided commitment had two clean instances and one partial, against a threshold of three used elsewhere. The node itself states which instance is weaker, rather than the count being rounded up to meet the rule.

And a glossary collision was found by accident. Two entries normalised to the same deduplication key and one was silently dropped, which was noticed only because a verification step checked for the term rather than trusting the insertion. The bug had been live for an unknown number of prior entries. An audit run at the close of this territory found exactly one collision across 1,277 entries, the one already fixed by hand, and no silent losses elsewhere, which is a better result than the failure deserved.

None of these changed a published claim. They are process failures, and the reason for listing them is that a territory concluding that measurement discipline is the operative variable should demonstrate some.

The counter-argument

Nine subjects chosen by one corpus is not a survey of enterprise AI. They were selected partly for having checkable numbers and partly for being interesting, and a territory assembled that way will find the patterns its assembler recognises. The organisational-constraint finding across four territories may be a house style rather than a regularity, and this corpus has said so before without changing its selection.

The commissioned-evidence complaint proves too much. Nearly all applied research in every field is funded by someone with an interest, medicine included, and the response there was disclosure and replication rather than dismissal. This territory has the disclosure and not the replication, which is a weaker position than it presents.

Naming missing studies is cheap. Any analysis can end by listing research nobody has done, and doing so converts an absence of evidence into an apparent finding. The five missing studies here are genuinely cheap and their absence genuinely tells you something, and that argument would sound identical if it were wrong.

And the self-report finding may be a selection artefact. Three instances where perception and measurement diverge were noticed because they diverge. Subjects where the two agree produce no article, and this corpus would not have found them.

The short version

Nine subjects, and the constraint was organisational in every one where anyone measured. A survey of 650 leaders put four of five failure causes outside capability. An audit of 47 engagements found the model almost never the blocker. Reliability turns on scaffolding, security on granted capability, lock-in on prompts, shadow AI on the absence of an approved tool.

Which is the fourth consecutive territory to land there, after the incident record, physical robotics and the cost of cheap generation. The intuitive explanation has not been the operative variable in any subject this corpus has examined with numbers attached.

And the evidence base is almost entirely commissioned. Four reliable anchors across nine subjects: a peer-reviewed benchmark, developer system cards published against interest, IBM's breach study, and a legal instrument. Everything else was published by a party selling the remedy it measured, including the most-quoted statistic in the field, which is a self-described preliminary working paper.

In five subjects, two literatures describe the same organisations and never meet. Sanctioned deployment against actual usage. Perception against telemetry. Benchmark against production. Attempts against exposure. Readiness against outcome. Each reconciling study is cheap, uses data the organisation already holds, and has no commercial sponsor.

The self-report is more favourable every time: 20% faster estimated against 19% slower measured, 97% benefiting against 29% returning, 59% reporting better code against refactoring at 3.8%. And trust falls with proximity to the work, with 26% of individual contributors reporting declining trust against 9% of directors.

Finally, every subject has somebody absorbing a cost they did not choose: the reviewer, the security function, the maintainer, the advertiser, the buyer. A burden nobody chose is a burden nobody measures, which is why the constraint stays invisible until somebody stops carrying it.

Common questions

What is the central finding of this territory? That the binding constraint on enterprise AI deployment was organisational rather than technical in every subject where anyone measured a cause. A March 2026 survey of 650 enterprise technology leaders attributed 89% of pilot failures to five causes, four of which are organisational or infrastructural, and a consultancy audit of 47 engagements from 2022 to 2025 states that the model was almost never the blocker. The same pattern appears where nobody framed it as a cause: reliability turns on scaffolding, security on what capability an agent is granted, lock-in on prompts and evaluations, and shadow AI on the absence of an approved alternative.

Is that finding reliable? It is consistent and it rests on a weak evidence base, which the territory states rather than glosses. There are four reliable anchors across nine subjects: a peer-reviewed benchmark paper, developer system cards disclosing attack success rates against interest, IBM's breach study with a stated method and multi-year series, and the AI Act as a legal instrument. Everything else was published by parties selling the remedy for the problem they measured. The territory can describe what commissioning parties measured and cannot say what is true in the gap between their questions.

What are the missing studies? Five, all cheap, all using data organisations already hold. Total organisational AI value covering governed and ungoverned use together. Paired survey and telemetry on the same commits, asking how a change felt and tracking its churn. Pass^k measured on a production agent rather than a benchmark. Attempts per agent per day, which converts a per-attempt breach rate into an exposure. And an A/B comparison of two deployments in one organisation, one on a readied data estate and one not. None has a commercial sponsor, because in each case the party best placed to run it is the party who would look worse.

Why is self-report always more favourable? Because of partial visibility rather than dishonesty. Effort saved is felt immediately; cost that arrives later, lands on someone else or spreads across a system is felt by nobody. Developers estimated 20% faster against a measured 19% slower, 97% of executives report benefiting against 29% reporting organisational return, and 59% of developers report improved code quality while refactoring fell to 3.8%. A workforce survey adds a fourth angle: 26% of individual contributors report declining trust with heavy AI reliance, against 17% of managers and 9% of directors, so trust falls with proximity to the work.

Who bears the cost of these deployments? Consistently, somebody who did not choose them. Reviewers absorb the check on generated code, with senior engineers reporting 20 to 35% more review time. Security functions absorb shadow AI detection without having set policy or been asked whether an approved tool exists. Maintainers absorb churn. Advertisers pay above clean rates for inventory every quality metric called premium. Buyers absorb switching costs created by engineers solving real problems well. In none of these did the absorbing party decline, because declining usually means declining the job.

Does this mean enterprise AI is failing? No. Adoption is near-universal, delivery throughput gains are measured, and the most-quoted failure statistic measures a specific and demanding bar: measurable profit-and-loss impact within six months on a sample weighted toward the study's own lowest-return function. What the territory establishes is what limited these deployments, not what the technology can do.

How does this relate to the previous territories? It is the fourth consecutive territory to find that the intuitive explanation was the wrong one. The incident record found no documented failure fixed by a better model. Physical robotics found task shape predicting feasibility where difficulty did not. The cost of cheap generation found a cost ratio rather than an output-quality problem. And enterprise deployment finds an organisational constraint. Whether that is a real regularity or a house style is a live question, and the falsification test remains a subject where the intuitive explanation turns out to be correct.

What is the strongest objection to this synthesis? That nine subjects chosen by one corpus is not a survey of enterprise AI. They were selected partly for having checkable numbers and partly for being interesting, and a territory assembled that way will find the patterns its assembler recognises. A second objection is that naming missing studies is cheap: any analysis can end by listing research nobody has done, converting an absence of evidence into an apparent finding, and that argument would sound identical if it were wrong.

Sources

Primary documents only. Where a claim rests on a single report, the entry says so.

  1. The nine Territory 10 articles Artifipedia This is a synthesis and makes no new factual claims. Every figure appears in one of the nine articles with its own sourcing and caveats. Evidence quality is stated per row in the opening table because it varies more across this territory than any previous one: four reliable anchors against a majority published by parties selling the remedy they measured.
  2. Navigating AI in the Workplace: 2026 SHRM The one figure in this synthesis not drawn from the nine articles: 26% of individual contributors reporting declining trust with heavy AI reliance, against 17% of managers and 9% of directors and above.

Further reading

The primary literature behind the claims above, drawn from the concept entries this post links to, so a claim carries the same source here as it does there.

  • Liang et al. (2023), GPT detectors are biased against non-native English writers — an independent finding on a question the vendors measuring the same tools were not asking. :: https://arxiv.org/abs/2304.02819 Commissioned Framing
  • Wong et al. (2021), External Validation of a Widely Implemented Proprietary Sepsis Prediction Model — a validation nobody was commercially motivated to run, performed independently. :: https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2781307 Commissioned Framing
  • METR (2025), randomised trial of experienced developers on their own repositories — self-estimated 20% speedup against a measured 19% slowdown. :: https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/ Self-Report Gap
  • Stenberg (2026), The end of the curl bug-bounty — evaluation burden externalised onto an unpaid maintainer until the programme closed. :: https://daniel.haxx.se/blog/2026/01/26/the-end-of-the-curl-bug-bounty/ Cost Externality

Learn the concepts

← All posts