Home/Blog/Evaluation & evidence/Ten domains, and the specification moved in every one
Ten domains, and the specification moved in every oneAnd both of those are teleoperated or moving totes.100%domains with real deployment20%that redefined nothingOF TEN DOMAINS EXAMINEDAnd both of those are teleoperated or moving totes.
And both of those are teleoperated or moving totes.

Ten domains, and the specification moved in every one

Territory 7 closes. Physical automation succeeded wherever the task could be changed, and the one place it has not is where nothing was negotiable.

TL;DR. Ten domains, and one pattern. Physical automation succeeded in every case where some part of the specification could be changed, and in none where it could not. Industrial robots got an engineered workspace. Vehicles got a drawn and mapped domain. Row crops were bred for machines. Delivery drones deleted the landing. Vacuums widened the tolerance. Construction moved the work indoors. Surgical robots did not redefine anything and are teleoperated: 2.6 million procedures and no autonomy. Humanoids are the explicit refusal to redefine, and the best-documented deployment is seven units. Underneath all of it sits a corpus of about a million robot trajectories against trillions of tokens for language, because text existed already and trajectories have to be performed. The framework's own weakness is that it can describe anything after the fact, so this article ends with the test that would break it.

---

Status: synthesis. No new factual claims. Every figure appears in one of the ten Territory 7 articles with its own sourcing, and each is linked where used. Sourcing quality varies sharply across these domains and the table below says so.

---

The ten

DomainWhat changedResultEvidence quality
IndustrialWorkspace engineered4.66m units in serviceIFR, strong
VehiclesDomain drawn and mapped220.6m rider-only milesPeer-reviewed, strongest
SurgicalNothing. Teleoperated2.6m procedures, no autonomyPeer-reviewed, interested authorship
AgricultureCrops bred for machines10m acres weeded, no pickingCompany-reported
WarehouseWorkspace engineered750k robots, transport onlyContested
DronesLanding deleted100m+ autonomous milesCompany-reported
DomesticTolerance widened32.7m units, 38% on tasksAnalyst estimate
ConstructionWork relocated indoors0.03% of spend on siteIndustry report
LearningData pooled across bodies+50% success, ~1m trajectoriesPeer-reviewed
HumanoidsNothing, by design7 units documentedContested

Finding one: the specification moved every time

Five forms, and every successful domain used at least one.

Environment engineering rebuilt the workspace: fixtures, fixed lighting, known part geometry. Industrial and warehouse robotics, and construction by relocating the work to a factory where the strategy becomes available again.

Domain narrowing specified where and when: mapped service areas, cities chosen substantially for climate, remote assistance reachable. Autonomous vehicles.

Target standardisation changed the thing being acted on: uniform height, simultaneous ripening, mechanical tolerance. The processing tomato was developed alongside its harvester as a joint programme. The machine did not learn to handle the plant.

Sub-task deletion removed the hardest step: a parachute into a five-metre zone, and the aircraft never lands anywhere but home.

Tolerance widening accepted a worse result far more often: a vacuum that misses corners and runs daily.

Two domains redefined nothing. Surgical robotics is teleoperation, so the human supplies everything the machine cannot. Humanoids are the wager that redefinition is unnecessary, and their documented deployments are moving totes on flat floors in facilities already built around material handling robots. The accommodation happened anyway; the form factor changed.

Finding two: difficulty does not predict feasibility

The distribution of success does not match intuition, and consistently in the same direction.

Recognising a ripe strawberry is trivial for a person and unsolved for a machine. Navigating a field is hard for a person and straightforward with GPS. Assembling a car body is heavy precision work and is automated. Folding a shirt is what a child does and is not.

What separates them is not difficulty. It is whether the specification is negotiable, and three properties determine that.

Is the failure reversible? A missed weed is caught next pass; a bruised strawberry is unsellable permanently. The same 95% accuracy means different things.

Is the judgement binary or continuous? A weed is or is not the crop. Ripeness is a spectrum with asymmetric costs on either side.

And is the difficult part the deliverable? A drone can skip landing because arrival was a means. A laundry robot that does everything but fold has done nothing.

Finding three: the home and the building site block everything

Two environments defeated all five strategies, for the same structural reason in different forms.

A home cannot be rebuilt around a machine, barely narrows because every house differs, cannot be standardised because the items are whatever the household owns, and contains no sub-task to delete because in a chore the manipulation is the product. Only tolerance widening remained, which is why floors, lawns and pools have robots and folding does not.

A building site fails the same way and adds one more: tolerance cannot widen, because a wall that is not plumb is not a partially built wall. Tolerances are specified, inspected and legally enforceable. The industry's answer was to move the work into a factory, and prefabrication grows at roughly 18% a year while on-site robotics sits at 0.03% of spend.

These are the two environments where automation is most wanted and least present, and the coincidence is not a coincidence.

Finding four: the data explains all of it

About a million robot trajectories underpin the largest generalist policies, with over 85% of the pooled corpus from four robot arms. Language models train on trillions of tokens.

Text existed already. Every book and comment was written by someone for their own reasons and the model received it as a byproduct. A robot trajectory has never existed until a robot performs it, on hardware, in real time, once.

So the five accommodations are not failures of ambition. They are what is available given the corpus that exists. Redefinition is how you build something useful without the data that would make redefinition unnecessary, and it works: 4.66 million industrial units and 220 million driverless miles are not consolation prizes.

What would break this

The honest problem with this framework is that any environment can be described as somewhat engineered after the fact, which would make it a description rather than a claim. So here is the test, stated before the evidence.

A deployment counts as falsifying if all four hold. More than a hundred units in continuous commercial operation. Materially different tasks performed by the same unit without reconfiguration. A site the operator did not modify, with modifications disclosed if any. And a published intervention rate per hour.

If that appears and the framework still explains it by finding some accommodation, the framework is unfalsifiable and should be discarded.

The corresponding prediction: the next domains to automate will be those with a negotiable specification, not those with an easy task. Watch for tasks where failure is reversible, judgement is binary, and the hard part is a means rather than the deliverable. That predicts more warehouse and logistics automation, more agricultural weeding, more inspection, and continued absence in laundry, cooking, care and on-site construction, regardless of how capable models become.

What this does not show

That physical AI is stalled. Ten domains with real deployments, one with 220 million driverless miles and peer-reviewed safety data, is not stagnation.

That the sample is representative. These ten were selected for having numbers, which is not the same as being typical. Domains without published figures are absent by construction.

That the evidence is comparable. It runs from peer-reviewed crash analysis to company-reported acreage, and the table says which is which. Conclusions drawn across it inherit the weakest links.

And nothing about employment. Whether these deployments displace, augment or relocate work is a separate question this territory did not examine.

The counter-argument

The framework is close to unfalsifiable and the test above may not save it. Deciding what counts as an unmodified site is a judgement, and a motivated reader can find engineering anywhere. That is the strongest objection and the test is an attempt to answer it rather than a proof that it fails.

Redefinition may just be engineering. Every solved problem was reframed to be solvable. If the claim describes all of engineering it is true and uninformative, and the distinction offered here, that a specific sub-task or tolerance changed, may be too fine to hold.

Selection produced the pattern. Domains were chosen because they had published figures, and publishing figures correlates with having a bounded deployable product, which correlates with having redefined the task. The pattern may be an artefact of which domains generate statistics.

And the data argument may be temporary. Simulation, human video and deployed-fleet collection could each change the corpus within a few years, at which point the accommodations become historical rather than structural.

The short version

Ten domains, one pattern: physical automation succeeded wherever some part of the specification could be changed, and nowhere it could not.

Workspaces engineered for 4.66 million industrial units. A domain drawn and mapped for 220 million rider-only miles. Crops bred for machines, with the processing tomato developed alongside its harvester. A landing deleted, with the aircraft releasing by parachute into a five-metre zone and never touching down away from home. A tolerance widened, so 32.7 million vacuums clean daily and adequately rather than weekly and well. And work relocated indoors when a building site refused all five.

Two domains redefined nothing. Surgical robotics is teleoperation: 2.6 million procedures and the machine decides nothing. Humanoids are the refusal to redefine, and the best-documented deployment is seven units moving totes on flat floors in facilities already built around material handling robots.

Difficulty does not predict feasibility. What predicts it is whether failure is reversible, whether judgement is binary, and whether the hard part is a means or the deliverable. The home and the building site fail all three, which is why the two environments where automation is most wanted are the two where it is least present.

And underneath everything sits about a million robot trajectories, 85% from four arms, against trillions of tokens for language. Text existed already. Trajectories have to be performed, once, in real time, on hardware.

The framework's weakness is that it can describe anything afterwards, so here is the test: more than a hundred units, materially different tasks without reconfiguration, an unmodified site, and a published intervention rate. If that arrives and this framework still explains it away, discard the framework.

Common questions

What is the main finding of this territory? That physical automation succeeded in every domain where some part of the task specification could be changed, and in none where it could not. Five forms of change recur: engineering the workspace, narrowing the operating domain, standardising the target, deleting the hardest sub-task, and widening the acceptable tolerance. Every successful domain used at least one.

Which domains redefined nothing? Two. Surgical robotics is teleoperation, so the surgeon supplies in real time everything the machine cannot do, across 2.6 million procedures with no autonomous decision-making. Humanoid robotics is the explicit bet that redefinition is unnecessary, and its documented deployments involve moving totes on flat floors in facilities already designed around material handling robots, which means the accommodation happened anyway and only the form factor changed.

Why does task difficulty not predict what gets automated? Because the deciding factor is whether the specification is negotiable, not how hard the task is. Recognising a ripe strawberry is trivial for a person and unsolved for machines; navigating a field is hard for a person and straightforward with GPS. Three properties separate them: whether failure is reversible, whether the judgement is binary or continuous, and whether the difficult part is a means to the goal or the goal itself.

Why are homes and building sites so resistant? Because they defeat all five strategies. Neither can be rebuilt around a machine, both vary enough that narrowing barely helps, neither target can be standardised, and in both the difficult manipulation is the deliverable rather than a means. A building site adds a fifth barrier: tolerances are specified, inspected and legally enforceable, so widening them is not available either. These are the two environments where automation is most wanted and least present.

What role does training data play? It explains why accommodation is the norm rather than the exception. The largest generalist robot policies rest on roughly a million trajectories, with over 85% of the pooled corpus from four robot arms, against trillions of tokens for language models. Text and images existed already as byproducts of human activity; a robot trajectory has never existed until a robot performs it, in real time, once, on hardware. The five accommodations are how useful systems get built without the corpus that would make them unnecessary.

What would prove this framework wrong? A deployment meeting four conditions: more than a hundred units in continuous commercial operation, materially different tasks performed by the same unit without reconfiguration, a site the operator did not modify with any modifications disclosed, and a published intervention rate per hour. If such a deployment appears and the framework still explains it by locating some accommodation, then it describes everything and predicts nothing, and should be discarded.

What does the framework predict? That the next domains to automate will be those with negotiable specifications rather than easy tasks. Expect continued expansion in warehousing, logistics, agricultural weeding and inspection, where failure is reversible and judgement is binary. Expect continued absence in laundry, cooking, personal care and on-site construction, where it is not, largely regardless of how capable underlying models become.

How reliable is the evidence across these ten domains? It varies more than any single article conveys, which is why the table states quality per row. Autonomous vehicle safety is peer-reviewed against regulatory crash data and is the strongest. Surgical outcomes are peer-reviewed with interested co-authorship. Warehouse injury data is contested between parties who both have a stake. Agricultural acreage, drone deliveries and humanoid unit counts are company-reported or analyst estimates with no disclosure regime behind them. Any conclusion drawn across the whole set inherits the weakest links in it.

Sources

Primary documents only. Where a claim rests on a single report, the entry says so.

  1. The ten Territory 7 articles Artifipedia This is a synthesis and makes no new factual claims. Every figure appears in one of the ten articles with its own sourcing, which ranges from peer-reviewed crash analysis to company-reported acreage. The table in the article states evidence quality per row for that reason.

Learn the concepts

← All posts