Home/Blog/Evaluation & evidence/220 million miles, inside a boundary Waymo drew
220 million miles, inside a boundary Waymo drewAdjusted to the streets Waymo drove. Which is correct, and is the limit.100%human benchmark, adjusted10%Waymo, serious injuryAdjusted to the streets Waymo drove. Which is correct, and is the limit.
Adjusted to the streets Waymo drove. Which is correct, and is the limit.

220 million miles, inside a boundary Waymo drew

The strongest safety evidence in physical autonomy, and the methodology that makes it honest is also what limits what it can tell you.

TL;DR. Through March 2026 the Waymo Driver had logged 220.6 million rider-only miles with no human in the vehicle. Against an adjusted human benchmark it reports 90% fewer serious-injury-or-worse crashes, 81% fewer any-injury crashes and 92% fewer pedestrian injury crashes, and peer-reviewed analysis in Traffic Injury Prevention finds the serious-injury reduction statistically significant. This is the best safety evidence physical autonomy has produced, and it is carefully done. The methodology is also the point of this article. The human benchmark is weighted to the specific streets Waymo drove, proportional to miles, because comparing against whole counties would be unfair. That is the correct choice. It also means the comparison is a statement about performance inside a boundary Waymo drew, and says nothing about anywhere else.

---

Status: established, with the methodology stated. Primary sources: Waymo's Safety Impact hub, which reports against NHTSA Standing General Order data, and the peer-reviewed analyses by Kusano and colleagues in Traffic Injury Prevention at 7.1 million and 56.7 million rider-only miles. Figures are theirs.

---

The numbers are real

It is worth saying clearly before anything else. This is not a demonstration and it is not a pilot.

220.6 million rider-only miles through March 2026, driven commercially without a human behind the wheel, in Phoenix, San Francisco, Los Angeles and Austin.

Against the adjusted human benchmark, across 127 million rider-only miles through September 2025: 90% fewer serious-injury-or-worse crashes, 81% fewer any-injury crashes, 92% fewer pedestrian injury crashes.

The peer-reviewed analysis at 56.7 million miles found a statistically significant reduction in suspected-serious-injury-or-worse crashes when all locations were combined, and reported 181 fewer any-injury, 78 fewer airbag-deployment and 11 fewer serious-injury crashes than the benchmark predicted over the same distance.

And both serious-injury crashes involving a Waymo in that period were secondary crashes, meaning the Waymo was not involved in the initiating event.

No other physical autonomy programme has published anything comparable. Compared to the deployment claims in the robotics article, this is what evidence looks like.

The methodology is the finding

Here is how the human benchmark is built, and it is done properly.

Crash and mileage data from the counties Waymo operates in would produce what the authors call an unadjusted benchmark: the crash rate of the whole county, including roads and conditions Waymo never drives. So the benchmark is subset and weighted to only the area within those counties where the service actually drove, proportional to the miles driven there.

That is the right decision. Comparing a vehicle operating on selected urban streets against a county average including highways, rural roads and everything else would flatter it enormously. Adjusting removes that advantage.

And the same adjustment is what bounds the claim. The comparison answers: on these streets, in these conditions, does the Waymo Driver crash less than a human would? The answer is a well-evidenced yes.

It does not answer what happens on streets outside that boundary, because there is no data from outside it, by construction. The papers say so; the authors note the operational design domain has changed over time and does not necessarily include entire counties.

The operational design domain is environment engineering

The previous article argued that physical automation succeeds where the environment has been engineered to remove variation, and that the boundary is controlled against open rather than symbolic against physical.

A city street cannot be rebuilt around a vehicle. So the domain is narrowed instead.

Geographic: a defined service area, mapped in high detail in advance, expanded deliberately rather than encountered.

Environmental: cities chosen substantially for climate. Phoenix has weather that removes an entire category of perception problem.

Operational: remote assistance available, so a vehicle that cannot resolve a situation can request guidance rather than having to solve it.

This is the same strategy as a factory cell, applied where the floor cannot be moved. You cannot control the world, so you select which part of it to enter, map it beforehand, and keep a human reachable.

That is not a criticism. It is an extremely effective engineering approach and the results establish it works. It is a description of what the achievement is: not that autonomy solved open-ended driving, but that a sufficiently narrowed domain makes driving tractable, and the narrowing can be widened over time.

The open question is the shape of that widening. Whether each expansion costs roughly the same effort, or whether cost rises as the remaining territory gets harder, is the thing nobody outside the programme can currently see.

What is not published

Three numbers would materially change what an outsider can conclude.

Remote assistance frequency. How often a vehicle requests human guidance per hundred miles. This is the intervention rate question, and it separates autonomy from a highly capable supervised system. Waymo publishes extensive safety data and this is not part of it.

Domain exclusion detail. What the service refuses. Roads not entered, conditions that suspend operation, times it will not run. The boundary's shape is as informative as the performance inside it.

And mileage denominators by condition. The aggregate is 220.6 million miles. The distribution across weather, night driving, road type and city is what would show whether the harder conditions are represented or avoided.

None of these are hidden in a suspicious sense. No regulation requires them, no competitor publishes them, and Waymo already discloses far more than anyone else. The point is that the transparency, which is genuine, is transparency about outcomes rather than about scope.

One number worth noticing

Analysis of California regulatory filings found the deadheading rate, the share of miles driven with no passenger aboard, improved from 51.5% in January 2024 to 44.3% by September 2025, with only about 54% of total California miles carrying a passenger across the study period.

Nearly half the miles are empty. That is a fleet logistics problem rather than a technology one, and it is a substantial cost that safety statistics do not touch. It is also a reminder that the operational question and the capability question are separate, and that the second being answered does not settle the first.

What this establishes

That physical autonomy can outperform humans on measured safety inside a defined domain. With peer review, regulatory crash data and a benchmark adjusted against its own interest. That is a genuine result and it is the strongest in the field.

That the domain is the mechanism. Not a limitation to be apologised for, but the thing that makes the performance achievable, and the reason the result cannot be extrapolated beyond it.

That published safety data is not the same as published scope data. An organisation can be exemplary on the first while the second remains invisible, and both are needed to know what a deployment means.

And that this is what the standard should look like. Article 128 asked for deployment counts, task breadth, environment disclosure and intervention rates. Waymo supplies the first at scale and the fourth not at all. That is still three quarters more than any comparable programme.

What is unresolved

Whether the safety advantage holds as the domain widens. Every expansion adds conditions the system has less experience of, and the current numbers are an average over a domain chosen partly for tractability.

How often remote assistance is used. Unpublished, and the single most informative missing number.

Whether the benchmark adjustment is right. Weighting to Waymo's own miles is defensible and it is also a choice with alternatives, and reasonable analysts differ on how to construct a fair comparison population.

And what underreporting does to the comparison. The authors themselves note that human crash underreporting is estimated from multiple sources, that confidence intervals on it are not straightforward, and that it may vary by locality. Any comparison against police-reported human data inherits that uncertainty.

The counter-argument

Asking for scope data understates a genuine achievement. Waymo publishes more than any comparable programme, submits to peer review, and adjusts its own benchmark in a direction that reduces its apparent advantage. Treating that as partial transparency risks penalising the one organisation doing it.

Every deployed system has an operational envelope. Aircraft autopilots, medical devices and industrial robots all specify conditions of use, and nobody treats that as a caveat undermining their performance. Singling out an operational design domain as though it were a concession applies a standard nothing else meets.

The domain has widened substantially. From a small area around Chandler in 2020 to most of San Francisco and multiple metropolitan areas, which is evidence about the widening question rather than silence on it.

And the intervention-rate framing may not transfer. Remote assistance in Waymo's architecture is not a person driving; it is a system requesting guidance on an ambiguous situation while remaining responsible for control. Counting those as interventions would compare unlike things.

The short version

220.6 million rider-only miles through March 2026, with 90% fewer serious-injury-or-worse crashes, 81% fewer any-injury crashes and 92% fewer pedestrian injury crashes against an adjusted human benchmark, and a statistically significant serious-injury reduction in peer-reviewed analysis at 56.7 million miles. Both serious-injury crashes in that period were secondary, with the Waymo not in the initiating event.

This is the best safety evidence physical autonomy has produced.

And the benchmark is weighted to the specific streets Waymo drove, because comparing against whole counties would flatter it. That adjustment is correct, and it means the result is a statement about performance inside a boundary Waymo drew.

The boundary is the mechanism. A street cannot be rebuilt around a vehicle, so the domain is narrowed instead: mapped service areas, cities chosen substantially for climate, remote assistance reachable. The same strategy as a factory cell, applied where the floor cannot be moved.

Three numbers would change what an outsider can conclude, and none is published: remote assistance frequency per hundred miles, what the domain excludes, and the mileage distribution across weather, night and road type. The transparency is real and it is transparency about outcomes rather than scope.

And nearly half the miles carry no passenger. Deadheading fell from 51.5% to 44.3%, with about 54% of California miles carrying a rider. The capability question being answered does not settle the operational one.

Common questions

How many miles has Waymo driven without a human driver? 220.6 million rider-only miles through March 2026, meaning miles driven with no human behind the wheel, in commercial service across Phoenix, San Francisco, Los Angeles and Austin. Rider-only is the relevant figure because it excludes miles with a safety driver present.

What do the safety numbers actually say? Measured against an adjusted human benchmark across 127 million rider-only miles through September 2025, Waymo reports 90% fewer serious-injury-or-worse crashes, 81% fewer any-injury crashes and 92% fewer pedestrian injury crashes. Peer-reviewed analysis at 56.7 million miles found the serious-injury reduction statistically significant when locations were combined, and both serious-injury crashes involving a Waymo in that period were secondary crashes where the Waymo was not part of the initiating event.

What does "adjusted benchmark" mean and why does it matter? The human comparison is not the crash rate of the whole county. It is subset and weighted to only the area within those counties where Waymo actually drove, proportional to miles driven there. That is the correct choice, because comparing selected urban streets against a county average including highways and rural roads would flatter the system substantially. It also bounds the conclusion: the comparison establishes performance on those streets and provides no evidence about anywhere else.

Is the operational design domain a weakness? No, it is the mechanism. A city street cannot be rebuilt around a vehicle the way a factory floor is rebuilt around a robot, so the domain is narrowed instead: geographic service areas mapped in advance, cities chosen substantially for climate, remote assistance available for situations the vehicle cannot resolve. That is an effective strategy and the results show it works. It means the achievement is that a sufficiently narrowed domain makes driving tractable, not that open-ended driving has been solved.

What is not published? Three things that would materially change outside analysis. How often a vehicle requests remote assistance per hundred miles, which is the number separating autonomy from highly capable supervision. What the domain excludes, since the shape of the boundary is as informative as performance inside it. And how the miles distribute across weather, night driving and road type, which would show whether harder conditions are represented or avoided.

Why does deadheading matter? Because it is a large operational cost that safety statistics do not touch. Analysis of California regulatory filings found the share of miles driven without a passenger improved from 51.5% in January 2024 to 44.3% by September 2025, with roughly 54% of total California miles carrying a rider across the period. Nearly half the miles are empty, which is a fleet logistics problem rather than a technology one, and it is a reminder that answering the capability question does not settle the operational one.

Does this prove autonomous vehicles are safer than humans? It provides strong evidence that this system, on these streets, in these conditions, has a lower crash rate than the adjusted human benchmark for the same streets. That is a meaningful and carefully evidenced claim. It is not a general claim about autonomous vehicles, or about driving outside the operating domain, and the authors are explicit about the limits, including that estimates of human crash underreporting carry uncertainty that the comparison inherits.

How does this fit the argument about controlled versus open environments? It supports it and refines it. Physical automation succeeds where variation is removed in advance. An industrial robot removes it by rebuilding the workspace; a self-driving service removes it by selecting which workspace to enter, mapping it beforehand and keeping a human reachable. Both are environment engineering. The interesting question is whether widening the domain costs a constant amount per expansion or an increasing one, and that is not visible from outside.

Sources

Primary documents only. Where a claim rests on a single report, the entry says so.

  1. Waymo Safety Impact Waymo, reporting against NHTSA Standing General Order data The 220.6 million rider-only miles through March 2026 and the crash reductions against the adjusted human benchmark.
  2. Comparison of Waymo Rider-Only crash rates by crash type to human benchmarks at 56.7 million miles Kusano et al., Traffic Injury Prevention, 2025 The peer-reviewed analysis, including the benchmark construction weighted to the area actually driven, and the finding that both serious-injury crashes were secondary.
  3. Comparison of Waymo rider-only crash data to human benchmarks at 7.1 million miles Kusano et al., Traffic Injury Prevention, 2024 The earlier analysis, and the discussion of human crash underreporting whose uncertainty any such comparison inherits.

Learn the concepts

← All posts