Home/Blog/Evaluation & evidence/The good numbers all came from obligations

The good numbers all came from obligations

Territory 8 closes. Across ten subjects the reliability of a figure tracked whether someone was legally required to produce it, and nothing else.

TL;DR. Ten subjects, and one predictor of whether a number could be trusted: was anyone obliged to publish it. Securities filings gave a write-down to the dollar. A statutory reporting regime gave data centre water use. Regulatory crash reporting gave vehicle safety. Everywhere else the best available figure came from a party with an interest, and the alternative was no figure at all. In six of the ten subjects the two most-quoted numbers were both accurate and measuring different things: per-query against total energy, run rate against booked revenue, direct water against water including generation, aggregate employment against a cohort, compute production against token usage. The territory's finding is not that AI's footprint is large or small. It is that the measurement layer fails in a specific and predictable way, and the failure is almost never dishonesty.

---

Status: synthesis. No new factual claims. Every figure appears in one of the ten Territory 8 articles with its own sourcing and its own caveats, and each is linked where used.

---

The ten

SubjectQuotedWhat answers the questionSource quality
ElectricityEnergy per queryComposition of the other 98%IEA, strong
DepreciationReported earningsThe useful-life assumptionSEC filings, strong
LithographyCountry riskOne supplier, any geographyCompany reporting
RevenueRun rateBooked revenue, stated basisPrivate, unauditable
Inference prices10x per year9x to 900x, by milestoneEpoch, strong
Water500 ml per response10 to 50 ml, and the basinStudy plus disclosure
LabourNo disruptionA 20% cohort holeMixed, both strong
FinancingDeal totalsA $3.5bn guarantee bookFilings plus reporting
GridChip supplyA four-to-seven-year queueSolid data, interested framing
Export controlsCompute shareUsage shareAdvocacy on both sides

Finding one: obligation predicts quality

Sort the ten by how much a figure can be relied on, and the sort is by disclosure regime.

The strongest numbers in this territory are filings. A write-down reported to the dollar because misreporting is an offence. A lease-guarantee cap of $3.5 billion because a 10-Q requires it. Vehicle crash data under a Standing General Order. Data centre water use under a directive requiring 24 indicators from facilities above 500 kW.

The weakest are the ones nobody must produce. AI revenue for private companies. Humanoid deployment counts. Agricultural acreage. Drone deliveries. Chinese token usage share.

And the middle is where a strong research body chose to measure something nobody required, as with the IEA on energy, Epoch on inference prices, or a university group on sepsis.

The correlation is not with importance. Whether AI displaces workers matters more than a quarterly write-down, and is measured worse. Quality tracks obligation, and obligation was set by legislatures worrying about something else entirely, mostly investor protection and vehicle safety.

Finding two: both numbers were usually right

In six of the ten subjects, the two figures in circulation were both accurate and answered different questions.

Energy per query and total sector consumption. The first is about 2% of the second, and the argument is conducted almost entirely in the first.

Run rate and booked revenue. 63% apart for one company in one year, both correct.

Direct cooling water and water including electricity generation. Roughly a thousandfold apart, both defensible, and the circulated version compounded that with a dropped divisor.

Aggregate employment and a cohort within one occupation. No economy-wide disruption and a fifth of an entry-level cohort gone, simultaneously.

Nine times per year and nine hundred, from one study on one day.

Compute production share and token usage share, moving in opposite directions in the same year.

Almost none of this is dishonesty. It is scope, basis and denominator choices that do not travel with the figures they qualify, and the qualifier is the first thing lost when a number is repeated.

Finding three: the constraint was rarely the one being discussed

Three subjects turned on identifying what actually limits the thing.

Chips were the constraint and stopped being one. Advanced packaging capacity expanded, shipments scaled, and a four-to-seven-year utility queue became binding instead while coverage continued to track allocation.

Country risk is the discussed lithography constraint and one supplier is the actual one, unchanged by where a fab stands.

And capability is almost never the constraint in physical automation, which Territory 7 established and which recurs here: money cannot buy a transformer that has not been manufactured.

The pattern is that attention follows what is expensive, novel or contested, and scarcity sits elsewhere.

Finding four: interested parties produced most of the good work

Worth stating because it cuts against an instinct.

The most precise per-query energy and water figures were published by an operator measuring its own product. The most comprehensive robotic-surgery synthesis was co-authored by the manufacturer. The best-organised revenue tracker is maintained by a research organisation with a view. The clearest export-control figures come from advocacy on both sides.

In none of these cases was the alternative a disinterested measurement. The alternative was nothing.

Which means the useful discipline is not to discount interested sources but to state the interest and read accordingly, and to notice when a party with an interest publishes a figure that cuts against it, as when a company's own reporting shows absolute consumption rising 27% while it reports a 12% emissions reduction.

What the territory does not show

That AI's physical footprint is small. Data centre electricity roughly doubling to around 945 TWh by 2030 is a substantial planning problem, and Ireland at 21% of national electricity is a constraint now.

That it is catastrophic. Just under 3% of global electricity by 2030, with per-unit efficiency improving throughout, is not the figure the loudest claims describe.

That the subjects were representative. These ten were chosen for having numbers, which selects for domains with disclosure regimes or motivated researchers, and systematically excludes subjects where nobody has measured anything.

And nothing about what any of it should mean for policy. The corpus does not take positions on contested political questions, and several of these are.

What would improve every one of them

A stated basis. Which scope, which population, which period, which denominator. It costs a clause and would have prevented most of the confusion documented across ten articles.

Disclosure obligations attached to the questions that matter rather than only to the ones that historically produced regulation. The EU's data centre reporting requirement is the clearest example of a rule producing checkable numbers where none existed.

And published sensitivity where a figure rests on a judgement. Two audited companies reached opposite conclusions about identical hardware, and neither filing tells a reader how much the answer moves within the defensible range.

The counter-argument

The obligation finding may be circular. Subjects with disclosure regimes were selected partly because they had figures worth writing about, so of course the figures are better. A territory chosen for its numbers will find that regulated numbers are good, which proves less than it appears to.

Stating a basis does not fix incentives. A party choosing a narrow scope and stating it clearly has still chosen the scope that flatters it, and disclosure of method is not the same as neutrality of method.

Six of ten is not a pattern. Four subjects turned on something else entirely, and describing the territory as being about measurement risks fitting a frame to a set of articles that were selected independently.

And the framing may understate genuine disagreement. Treating conflicts as definitional is comfortable and sometimes wrong: on labour, on export controls and on financing, there are people who disagree about the world and not merely about what to count. Reducing those to scope problems is its own kind of evasion.

The short version

Ten subjects, and one predictor of whether a figure could be trusted: whether anyone was obliged to publish it.

The strongest numbers were filings. A write-down to the dollar. A $3.5 billion guarantee cap in a 10-Q. Crash data under a standing order. Water use under a directive covering facilities above 500 kW. The weakest were the ones nobody must produce: private-company revenue, humanoid deployments, acreage, token share. The correlation is with obligation, not importance, and the obligations were written by legislatures worrying about investor protection and road safety.

In six of ten, the two circulating figures were both accurate and answered different questions. Energy per query against sector total, where the first is 2% of the second. Run rate against booked revenue, 63% apart. Direct water against water including generation, a thousandfold apart. Aggregate employment against a cohort. Nine times a year against nine hundred, from one study. Compute share against usage share.

Almost none of it was dishonesty. It was scope, basis and denominator choices that did not travel with the numbers, and the qualifier is the first thing lost in repetition.

In three, the binding constraint was not the one being discussed. Chips stopped binding and a utility queue started. Country risk is discussed and one supplier is the constraint. Capability is discussed and tolerance, or a transformer, is the limit.

And most of the good work came from interested parties, because the alternative was not a disinterested measurement but no measurement. The discipline that follows is to state the interest and read accordingly, not to discount the source.

Common questions

What is the main finding of this territory? That the reliability of a figure tracked whether someone was legally obliged to publish it, and very little else. Securities filings, regulatory crash reporting and statutory data centre disclosure produced numbers that can be checked. Everywhere else the best available figure came from a party with an interest, and the alternative was no figure at all. The correlation is with obligation rather than with how much the question matters.

Why were so many disputes definitional rather than factual? Because scope, basis and denominator choices change what a number measures without changing its precision, and those qualifiers are the first thing lost when a figure is repeated. In six of the ten subjects, the two most-quoted numbers were both accurate: energy per query against total sector consumption, run rate against booked revenue, direct water against water including electricity generation, aggregate employment against a cohort, one price-decline rate against another, and compute production against token usage.

Does this mean the numbers are all unreliable? No. It means their reliability varies systematically and predictably, and the variation is visible in advance. A figure from a filing subject to legal penalty for misstatement is different in kind from a company-reported acreage figure with no disclosure regime behind it. The useful habit is to check which kind you are looking at before deciding how much weight it carries.

Should interested sources be discounted? Not automatically, because in most of these subjects the alternative was nothing. The most precise per-query energy and water measurements were published by an operator measuring its own product, and the most comprehensive surgical outcomes synthesis was co-authored by a manufacturer. The discipline is to state the interest and read accordingly. It is also worth noticing when an interested party publishes something that cuts against itself, as when a company reports a 12% emissions reduction alongside 27% growth in absolute consumption.

What was the most common structural error? Substituting a per-unit measure for a total, or an aggregate for a cohort. Energy per query is about 2% of AI data centre consumption and is close to the whole of the public argument. Aggregate employment stability is routinely reported as evidence that nobody was affected, which it is not. Both are the same mistake in opposite directions.

Why did the binding constraint keep being the wrong one? Because attention follows what is expensive, novel or contested, and scarcity is often none of those. Chip allocation was the constraint two years ago and coverage still tracks it, while a four-to-seven-year utility interconnection queue now determines what gets built. Lithography concentration is discussed as a country risk when the actual dependency is one supplier regardless of geography.

What single change would improve most of this? A stated basis attached to every figure: which scope, which population, which period, which denominator. It costs a clause and would have prevented most of the confusion documented across these ten articles. Beyond that, disclosure obligations attached to questions that matter rather than only to those that historically produced regulation, and published sensitivity ranges where a number rests on a judgement.

What is the strongest objection to this synthesis? That the obligation finding is partly circular. These subjects were chosen because they had numbers worth examining, which selects for domains with disclosure regimes or motivated researchers, so finding that regulated numbers are better proves less than it appears to. A second objection is that treating conflicts as definitional can be evasive: on labour, export controls and financing there are people who disagree about the world and not merely about what to count.

Sources

Primary documents only. Where a claim rests on a single report, the entry says so.

  1. The ten Territory 8 articles Artifipedia This is a synthesis and makes no new factual claims. Every figure appears in one of the ten articles with its own sourcing and caveats, ranging from securities filings and statutory returns to advocacy estimates, and the table states source quality per row for that reason.

Further reading

The primary literature behind the claims above, drawn from the concept entries this post links to, so a claim carries the same source here as it does there.

  • Raji et al. (2021), AI and the Everything in the Whole Wide World Benchmark — operationalisation determining what a measurement can support. :: https://arxiv.org/abs/2111.15366 Scope Boundary
  • Wong et al. (2021), External Validation of a Widely Implemented Proprietary Sepsis Prediction Model — the difference between overall performance and performance on the population that matters. :: https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2781307 Scope Boundary

Learn the concepts

← All posts