Home/Blog/Evaluation & evidence/32 million vacuums, and 38% on household tasks
32 million vacuums, and 38% on household tasksThe vacuum works because floors forgive. Folding does not.100%tasks in the benchmark38%completedOF 1,000 HOUSEHOLD TASKSThe vacuum works because floors forgive. Folding does not.
The vacuum works because floors forgive. Folding does not.

32 million vacuums, and 38% on household tasks

The most-deployed consumer robot succeeded by accepting a worse result more often. The home is the one place where the other ways of making a task tractable are unavailable.

TL;DR. Cleaning robots shipped 32.72 million units in 2025 on IDC estimates, up 20.1% year on year, making the robot vacuum the most-deployed consumer robot in history. Meanwhile Stanford's 2025 BEHAVIOR benchmark showed 38% completion across 1,000 household tasks, and China committed CNY 10 billion in the same year to humanoid research with a household focus. The vacuum did not solve floor cleaning. It changed the standard. It cleans less thoroughly than a person and cleans far more often, unattended, which is a better trade for floors and a worse one for almost everything else. And the home is the single environment where the other four ways of making a task tractable are all unavailable, which is why the vacuum is forty years old and the laundry robot is not here.

---

Status: established, with market figures attributed. Shipment figures are IDC estimates reported in industry coverage. The BEHAVIOR benchmark result is from Stanford's 2025 release. Household robotics has no regulatory disclosure regime, so commercial figures are analyst and company estimates rather than audited counts.

---

The most-deployed consumer robot there is

32.72 million cleaning robots shipped in 2025, up 20.1% on the previous year, with smart vacuums the largest segment. Three Chinese suppliers took 62% of shipments, which is the signature of a commoditised category rather than an emerging one.

Vacuuming and mopping accounted for roughly a third of the household robot market by value. Obstacle avoidance in current models is reported above 95% accuracy in real-world tests, and mid-range units now carry lidar, mapping and self-emptying docks.

By unit count this dwarfs every other robot in this territory. More cleaning robots ship in a year than the entire installed base of industrial robots accumulated in a decade.

It is not very good at cleaning

Also true, and both facts matter.

Battery life averages 90 to 120 minutes, which industry analysis describes as insufficient for larger homes without manual intervention. Complex layouts still cause incomplete cleaning cycles. Anyone who owns one knows it misses corners, gets trapped, and requires rescuing.

A person with an upright vacuum cleans a floor better in less elapsed time.

So why did it win?

Because it changed the standard

The task was not "clean the floor as well as a person." It became "keep accumulated dust below the level where the floor looks dirty, without anyone doing anything."

Those are different tasks and the second is much easier.

Frequency substitutes for thoroughness. A person vacuums once a week and does it well. A robot vacuums daily and does it adequately. Integrated over a week the floor is cleaner, and the human cost went to zero.

Unattended operation is the whole product. The value is not the cleaning quality; it is that nobody was present. That is why the comparison to a human with a vacuum is the wrong comparison, and why the market grew 20% in a year on a device that misses corners.

This is a fifth form of task redefinition, alongside the four the territory has already found: environment engineering, domain narrowing, target standardisation and sub-task deletion. Call it tolerance widening: accept a worse result more often, where more often is worth more than better.

And the home blocks the other four

Here is why the vacuum has no siblings after forty years of trying.

Environment engineering is unavailable. You cannot rebuild someone's home around a machine. The factory strategy is off the table by definition.

Domain narrowing barely helps. A robot could be specified for one house, and every house differs in layout, flooring, furniture, clutter, pets and lighting. Unlike a mapped urban service area, the population of homes has no shared structure to exploit.

Target standardisation is impossible. Row crops were bred for machines. Laundry cannot be. The items are whatever the household already owns, in whatever condition, and nobody is replacing their wardrobe to suit a robot.

Sub-task deletion has nothing to delete. A delivery drone can drop by parachute because arrival was a means, not the end. In a household chore the difficult manipulation is the product. A laundry robot that does everything except the folding has done nothing.

Which leaves tolerance widening as the only available move, and it only works where frequency beats quality.

Which tasks does that fit?

The test is whether an adequate result repeated often beats a good result done rarely.

Floors: yes. Dust accumulates continuously, partial removal is genuinely useful, and there is no failure state. Missing a corner today is fixed tomorrow.

Lawns: yes, for the same reasons, which is why robot mowers are the second consumer category.

Pools: yes. Same structure again.

Folding laundry: no. A badly folded shirt is not partially folded; it has to be redone. The failure is not partial, and doing it more often does not help, which is irreversible failure in a domestic setting.

Loading a dishwasher: no. A wrongly loaded item does not get cleaned, and a broken glass is worse than an unloaded one.

Cooking: no. Every step gates the next and errors compound rather than average out.

The pattern is clean. Tolerance widening works on tasks that are continuous, partially completable and forgiving. It fails on tasks that are discrete, all-or-nothing and unforgiving. And almost every household chore people actually want automated is in the second category.

What the benchmark says

Stanford's 2025 BEHAVIOR benchmark reported 38% completion across 1,000 household tasks.

That figure is worth sitting with, because it is measured rather than asserted, and because it is the number the humanoid case has to move.

China committed CNY 10 billion, about $1.4 billion, to humanoid research in 2025 with a household focus. Japan widened long-term care subsidies to cover half of eligible robot costs in private residences. Language-model-based task planning APIs for robotics arrived the same year.

The investment is real and the target is the hardest environment available, chosen precisely because it is where the demand is.

Three things this establishes

Tolerance widening is a real strategy and it is narrow. It produced the most-deployed consumer robot in history and it applies to a small set of task shapes: continuous, partially completable, forgiving. Recognising which shape a task has predicts feasibility better than assessing how hard it looks.

The home is the adversarial case for automation, not the easy one. Intuition says a factory is industrial and difficult while a house is domestic and simple. The opposite is true structurally, because a factory can be rebuilt and a house cannot, and the discussion consistently gets this backwards.

And 38% on household tasks is the number to watch. It is a measured baseline on a defined task set, which means progress against it is checkable in a way that demonstration videos are not.

What it does not establish

That household robots will not work. Substantial capital and state funding are behind the attempt, and 38% is a baseline rather than a limit.

That the shipment figures are audited. They are analyst estimates, and household robotics has no disclosure regime.

That BEHAVIOR measures the right thing. It is a simulation benchmark with a defined task set, and construct validity applies to it as much as to anything else. A score on it is evidence about it.

And that the vacuum's success was accidental. Redefining the task was a deliberate and clever product decision, not a consolation prize, and reading it as a failure to solve cleaning misreads what was built.

What is unresolved

Whether learned manipulation moves the 38%. Large-scale policy learning is the current bet and its trajectory on this task set is not established.

Whether any chore other than the three current ones fits tolerance widening. Nobody has found a fourth, which is either a gap in imagination or evidence the set is small.

Whether homes get standardised in the other direction. Appliances designed to be robot-operable, laundry designed to be machine-foldable, are conceivable, and that would be target standardisation arriving late.

And what the intervention rate is. Owners rescue their vacuums regularly and nobody publishes how often, which is the same missing number as everywhere else in this territory.

The counter-argument

Calling the vacuum a lowered standard is unfair to the engineering. Simultaneous localisation and mapping in a cluttered dynamic environment, on a consumer battery and price point, is a genuine achievement, and above 95% obstacle avoidance is not a compromise.

The task-shape argument may be post hoc. Floors, lawns and pools succeeded, so their shared properties get identified as the reason. Whether those properties predict the next success has not been tested, and a framework that only explains what already happened is a description.

Tolerance widening may extend further than argued. A laundry robot that folds imperfectly but folds everything might be acceptable to many households, and the assumption that folding is all-or-nothing is an assertion about preferences rather than physics.

And the home may be narrowable after all. Purpose-built assisted-living residences, standardised furniture, designated robot zones. Care settings in Japan are moving this way, and that is domain narrowing which this article said was unavailable.

The short version

32.72 million cleaning robots shipped in 2025, up 20.1%, with three suppliers taking 62% of shipments. By unit count it is the most-deployed robot category anywhere, and more ship in a year than industrial robots accumulated in a decade.

It is also not very good at cleaning. Battery life of 90 to 120 minutes is insufficient for larger homes without intervention, complex layouts still cause incomplete cycles, and a person with an upright vacuum does a better job faster.

It won by changing the standard. Not "clean as well as a person" but "keep dust below visible without anyone doing anything." Frequency substituted for thoroughness, and unattended operation was the entire product. That is a fifth form of task redefinition alongside environment engineering, domain narrowing, target standardisation and sub-task deletion: tolerance widening.

And the home blocks all four of the others. You cannot rebuild someone's house. Every house differs, so narrowing barely helps. Laundry cannot be bred for machines the way row crops were. And in a chore the difficult manipulation is the product, so there is no sub-task to delete.

Which leaves tolerance widening, and it fits a narrow set of task shapes: continuous, partially completable, forgiving. Floors, lawns and pools qualify. Folding laundry does not, because a badly folded shirt has to be redone, and doing it more often does not help. Almost every chore people want automated is in the second category.

Stanford's BEHAVIOR benchmark reported 38% completion across 1,000 household tasks in 2025, while China committed about $1.4 billion to humanoid research with a household focus. The investment is aimed at the hardest environment available, and it was chosen because that is where the demand is.

Common questions

How many robot vacuums are there? Cleaning robots shipped 32.72 million units in 2025 on IDC estimates, up 20.1% year on year, with smart vacuums the largest segment and three Chinese suppliers taking 62% of shipments. By unit count this is the most-deployed robot category anywhere: more ship in a single year than the entire installed base of industrial robots accumulated over a decade.

Are they actually good at cleaning? Not compared to a person. Battery life averages 90 to 120 minutes, which industry analysis describes as insufficient for larger homes without manual intervention, and complex layouts still produce incomplete cleaning cycles. Someone with an upright vacuum cleans a floor better in less elapsed time.

Then why did they succeed? Because the task changed. It is not "clean the floor as well as a person" but "keep accumulated dust below visible without anyone doing anything." Frequency substitutes for thoroughness: a robot cleaning daily and adequately leaves a cleaner floor over a week than a person cleaning weekly and well, and the human cost is zero. Unattended operation is the product, not the cleaning quality.

What is tolerance widening? Accepting a worse result more often, where more often is worth more than better. It is a fifth form of task redefinition alongside environment engineering, domain narrowing, target standardisation and sub-task deletion, and it is the only one of the five available in a home.

Why is the home so hard for robots? Because the other four strategies are all blocked. You cannot rebuild someone's house around a machine, which rules out the factory approach. Every house differs in layout, flooring, clutter and lighting, so specifying a narrow domain barely helps. Laundry and dishes cannot be standardised the way row crops were bred for harvesters, since the items are whatever the household owns. And in a household chore the difficult manipulation is the product, so there is no sub-task to delete.

Which chores fit tolerance widening and which do not? Tasks that are continuous, partially completable and forgiving fit: floors, lawns, pools, which is exactly the set of consumer robot categories that exists. Tasks that are discrete, all-or-nothing and unforgiving do not: a badly folded shirt has to be redone rather than being partially folded, a wrongly loaded dish does not get cleaned, and cooking errors compound because each step gates the next. Almost every chore people most want automated is in the second group.

What does the BEHAVIOR benchmark show? Stanford's 2025 release reported 38% completion across 1,000 household tasks. It is a measured baseline on a defined task set rather than an assertion, which makes progress against it checkable in a way that demonstration videos are not. As with any benchmark, a score on it is evidence about it, and construct validity applies.

Is anyone likely to solve this? It is the most heavily funded target in robotics. China committed roughly $1.4 billion in 2025 to humanoid research with a household focus, Japan widened long-term care subsidies to cover half of eligible robot costs in private residences, and language-based task planning for robots arrived the same year. The investment is real and aimed at the hardest environment available, chosen precisely because that is where the demand is. Whether learned manipulation moves the 38% is the open question and it is not settled either way.

Sources

Primary documents only. Where a claim rests on a single report, the entry says so.

  1. Household robots market analysis Mordor Intelligence, January 2026 The BEHAVIOR benchmark figure of 38% completion across 1,000 household tasks, the CNY 10 billion Chinese humanoid commitment, and the supplier concentration.
  2. Cleaning robot shipment estimates IDC, reported in industry coverage The 32.72 million units shipped in 2025 and 20.1% year-on-year growth. Analyst estimate rather than an audited count; household robotics has no disclosure regime.

Further reading

The primary literature behind the claims above, drawn from the concept entries this post links to, so a claim carries the same source here as it does there.

  • International Federation of Robotics, World Robotics 2025 — the installed base that environment engineering produced. :: https://ifr.org/worldrobotics/report-2025 Task Redefinition
  • Kusano et al. (2025), Comparison of Waymo Rider-Only crash rates by crash type to human benchmarks at 56.7 million miles — domain narrowing, and a benchmark correctly adjusted to it. :: https://waymo.com/research/comparison-of-waymo-rider-only-crash-rates-by-crash-type-to-human-benchmarks/ Task Redefinition
  • Raji et al. (2021), AI and the Everything in the Whole Wide World Benchmark — why general benchmarks cannot carry general claims. :: https://arxiv.org/abs/2111.15366 Construct Validity
  • Bowman & Dahl (2021), What Will it Take to Fix Benchmarking in Natural Language Understanding? — what a benchmark must satisfy to support inference. :: https://arxiv.org/abs/2104.02145 Construct Validity
  • Parasuraman & Riley (1997), Humans and Automation — the supervisory role and its failure modes. :: https://journals.sagepub.com/doi/10.1518/001872097778543886 Teleoperation

Learn the concepts

← All posts