Zillow Offers: $304 million, in the audited filing
A pricing model moved from advising consumers to committing capital. The write-down appears in a quarterly SEC filing, which makes this the best-documented AI failure in the record.
TL;DR. Zillow used a valuation model to decide which houses to buy and at what price. In its third-quarter 2021 filing the company recorded a $304.4 million inventory write-down, stating the cause plainly: it had bought homes above its own current estimates of what they would sell for. It expected a further $240 to $265 million in the following quarter, announced it would wind down the business, and cut about 25% of its workforce. This is the best-documented case in the record, because the loss is a line in an audited filing rather than an estimate by anyone outside the company. And the model was not obviously broken. The same error rate that is acceptable when advising a consumer is ruinous when committing capital, and nothing about the model changed when the use did.
---
Status: established. Primary sources: Zillow Group's Form 10-Q for the quarter ended 30 September 2021, and the company's earnings release of 2 November 2021, both filed with the SEC. The figures below are the company's own.
---
Zillow's valuation estimate had existed for years as a consumer feature: an indicative price attached to a listing, used by people deciding whether to sell and by buyers deciding whether to look.
Zillow Offers used the same underlying capability differently. The company bought houses directly, renovated them and resold them. The valuation was no longer advice. It was the basis on which the company committed its own money.
By the third quarter of 2021 that business represented 56% of total revenue, $2.6 billion of $4.2 billion.
Then the filing. A write-down of $304.4 million on homes inventory, attributed in the 10-Q to purchasing homes at prices higher than the company's current estimates of future selling prices. A further $240 to $265 million of losses expected in the fourth quarter, primarily on homes it was still contractually obliged to buy.
On 18 October the company stopped signing new contracts, citing renovation and operational capacity constraints. On 2 November it announced the wind-down and a workforce reduction of about 25%.
The chief executive's stated reason was that the unpredictability in forecasting home prices far exceeded what the company anticipated.
Why this case is worth more than the others
Every other entry in this record depends on someone outside the organisation characterising what happened. This one does not.
The number is in an audited quarterly filing, subject to securities law, reviewed by auditors, and stated by the company against its own interest. There is no reliance on anonymous sources, no dispute about whether harm occurred, and no question about the figure.
The cause is stated by the company in the same document. The 10-Q says the write-down was primarily due to purchasing homes at higher prices than current estimates of future selling prices. That is a model-error explanation given by the party that would most prefer a different one.
And the consequence is unambiguous. Around $8 billion of market value, roughly 2,000 jobs, and the closure of the segment that produced most of the company's revenue.
Against the standard this record uses, this is the strongest entry available: primary source, established causation, quantified harm, and the operator's own account is the account.
The mistake was not the model
This is the part almost every retelling gets wrong, and it is the transferable lesson.
A valuation estimate carrying a few percent of typical error is a good consumer product. A homeowner told their house is worth roughly a certain amount, give or take, has been given something useful. The error is understood, the stakes are informational, and the reader supplies their own judgement.
The same estimate carrying the same error, used to decide a purchase price, is a different instrument entirely. Now the error is not a caveat, it is a position. Buy ten thousand houses at prices drawn from a distribution centred slightly too high, in a market that stops rising, and the aggregate error is a balance-sheet event.
Nothing about the model had to get worse for this to happen. The tolerance for its error changed by a factor of a hundred when the output stopped being advice and started being a bid.
Two things compounded it. Scale: an error repeated across a large inventory does not average out if it is a bias rather than noise. And direction: the business bought when the model said a house was underpriced, which means it systematically selected for cases where the model was too high. A model that is unbiased on average is not unbiased on the subset it causes you to act on.
What made it worse rather than what caused it
Three operational factors appear in the filings and the surrounding analysis, and it is worth separating them from the model question.
Renovation capacity. The company cited renovation and operational constraints when pausing purchases. Homes bought and not resold sit on the balance sheet accruing carrying cost while the market moves.
Resale throughput. The 10-Q notes closings pushed from one quarter into the next because of resale capacity limits. Holding period is a risk multiplier: the longer between purchase and sale, the more the forecast has to be right about.
And market turn. Prices in many areas had risen sharply through 2020 and early 2021, and the pattern changed. A model fitted on a rising market extrapolates a rising market.
None of these is an AI failure. All of them determine how much an AI failure costs, which is the more useful framing: the model set the exposure, the operations set the duration, and the market set the outcome.
It was predicted
A detail that gets little attention and should get more.
Academic work published in December 2020, a year before the wind-down, had already described the structural difficulty. Research on residential real estate intermediation examined the trade-off facing anyone buying homes to resell: to attract sellers who want liquidity you must pay near market value, which leaves a thin margin against the risk you have absorbed.
The finding was not about machine learning. It was about the economics of the position. A model that prices perfectly still leaves you holding inventory in a market that can move, and the margin available for being wrong is narrow by construction.
Which means the failure was legible in advance to anyone reading the finance literature rather than the technology coverage. That is a recurring pattern: the discipline that already understands the problem is frequently not the one building the system.
What this establishes, narrowly
A model's acceptable error rate is a property of its use, not of the model. Moving an estimate from an advisory context to a transactional one changes the required accuracy by orders of magnitude without changing a line of code.
Selection on model output creates bias even in an unbiased model. If you act when the model says a thing is cheap, you act disproportionately on its overestimates.
And a model that commits capital is a trading strategy, subject to the risk management any trading strategy requires: position limits, holding-period assumptions, and a view on what happens when the regime changes.
What it does not establish
That the model was bad. No accuracy figure for the purchase-decision model has been published, and the consumer estimate's published error is not the same instrument. The company's own explanation is about forecasting unpredictability, which is a statement about the market as much as the model.
That AI caused the loss. The filing attributes it to buying above current resale estimates. Whether a human team pricing the same inventory would have done better is unknown, and the finance research suggests the position was difficult regardless of who priced it.
That the business was unviable. Competitors continued operating iBuying programmes after this. What the case establishes is that this operator, at this scale, in this market turn, could not make it work.
And nothing about consumer harm. The losses fell on the company, its shareholders and about 2,000 employees. No homeowner was defrauded; sellers received the prices they agreed to, and in many cases those prices were above what the market would later support.
The three cases so far, and what separates them
Three entries into the record, the evidence quality varies more than the failures do. Setting them side by side is the clearest argument for why the status line at the top of each one matters.
Moffatt v Air Canada. A published tribunal decision, named parties, a stated finding, an award of $650.88. Established, and procedurally slight: small claims, no counsel, contracts never filed.
Amazon's hiring tool. One news investigation, five anonymous sources, no technical report, no figures, no named institutions, and an operator disputing that anyone was affected. Reported, not established. A near-miss on the available evidence.
Zillow Offers. An audited quarterly filing with a stated figure and a stated cause, given by the company against its own interest. Established, quantified, and undisputed.
The three are cited at roughly the same confidence and they are not remotely the same claim.
Notice which is best documented. It is not the one where a regulator investigated, or the one a court decided. It is the one where the operator had a legal obligation to tell its shareholders. Securities disclosure produced better evidence about an AI failure than any AI-specific mechanism has.
That is worth sitting with, because every proposed remedy in this field is an AI-specific mechanism: incident registries, mandatory audits, model cards, transparency requirements. The best evidence in this record came from a disclosure regime written in the 1930s for an entirely different purpose, which works because it attaches to money rather than to technology, and because failing to comply is a securities offence rather than a compliance finding.
The practical implication for anyone assessing an AI failure: look for the financial filing first. If the failure cost enough to be material, it is described somewhere with an auditor's name attached, and that description will be more reliable than anything written about it since.
How to avoid this specific failure
Five questions, and they apply to any model whose output triggers a commitment.
What is the cost of being wrong by the model's typical error? Not its worst case, its ordinary one. Multiply by volume. If that number is uncomfortable, the model is not accurate enough for the use regardless of how it benchmarks.
Does acting on the output select for the errors? Buying when a model says underpriced selects for overestimates. Approving when a model says low-risk selects for underestimates. This is nearly universal in decision systems and nearly never modelled.
What is the exposure duration? Time between decision and outcome is time for the forecast to be wrong. A model whose predictions resolve in an hour and one whose predictions resolve in six months carry entirely different risk from the same error rate.
What happens when the regime changes? A model fitted on one market condition extrapolates it. The question is not whether that happens but how much it costs when it does.
And is there a position limit? The single control that would have bounded this is the one that has nothing to do with machine learning: a cap on how much inventory the model is allowed to acquire before someone reviews whether it is working.
What is unresolved
Whether the model was inaccurate or the market was unforecastable. The company's language points at the second. The distinction matters for whether better modelling would have helped, and no public evidence settles it.
How much of the loss was operational. Renovation and resale constraints appear in the filings as material factors. Nobody has separated the share attributable to pricing from the share attributable to holding inventory too long.
Whether a position limit was considered. No public record indicates whether anyone proposed capping purchase volume pending performance review, or whether that was raised and overruled.
And what the model's actual error was on purchased homes. The most useful number in the whole case has never been published.
The counter-argument
Calling this an AI failure is a category error. A company took a directional position in a housing market with thin margins and a long holding period, and the market moved. That is an inventory risk failure. The pricing model is where the exposure was set, and describing the outcome as an algorithm going wrong obscures that a human strategy chose to take the position at all.
The decision to stop was correct and is treated as the scandal. Recognising within a quarter that the business produced unacceptable balance-sheet volatility, disclosing it, and closing the largest revenue segment is decisive management. The same instinct that praises a company for cutting losses treats this one as a debacle because the losses were large and the technology was novel.
Competitors did not all fail. Other iBuying operations continued. If the technology were the problem, the failure would have been universal, and it was not, which points at execution, scale and timing rather than at pricing models in general.
And the error may not have been large. Buying above eventual resale value by a modest percentage across a large inventory produces a nine-figure write-down without the model being notably inaccurate. The size of the loss reflects the size of the position, and reading it as a measure of model quality confuses the two.
The short version
Zillow used a valuation model to decide which homes to buy and at what price. Its third-quarter 2021 filing recorded a $304.4 million inventory write-down, attributed in the document to purchasing homes above the company's own current estimates of future selling prices, with a further $240 to $265 million expected the following quarter. It wound down the business, which had been 56% of revenue, and cut about 25% of its workforce.
This is the strongest entry in the record because the number is a line in an audited filing. No anonymous sources, no dispute about whether harm occurred, and the causal explanation comes from the party that would most prefer a different one.
The mistake was not the model. The same estimate carrying the same error is a good consumer product and a ruinous basis for a bid. The tolerance for its error changed by orders of magnitude when the output stopped being advice and became a purchase price, and nothing in the model changed at all.
Two things compounded it. Scale, because a biased error repeated across a large inventory does not average out. And selection, because buying when the model says a house is underpriced means acting disproportionately on the cases where the model was too high. A model unbiased on average is not unbiased on the subset it causes you to act on.
Operational factors set the cost rather than the cause: renovation capacity, resale throughput, and a market that stopped rising. And the structural difficulty had been described in finance research published a year earlier, which is a recurring pattern worth noticing. The discipline that already understands a problem is frequently not the one building the system.
The control that would have bounded this has nothing to do with machine learning: a limit on how much the model is allowed to buy before someone checks whether it is working.
Common questions
What happened with Zillow Offers? Zillow bought homes directly, renovated them and resold them, using a valuation model to decide which to buy and at what price. In its filing for the quarter ended 30 September 2021 it recorded a $304.4 million inventory write-down, stating the cause as purchasing homes at prices above its own current estimates of future selling prices, and expected a further $240 to $265 million of losses the following quarter. It announced the wind-down on 2 November 2021 along with a workforce reduction of about 25%.
How much did Zillow lose on its home-buying algorithm? The company recorded $304.4 million in its third-quarter 2021 filing and expected $240 to $265 million more in the fourth quarter, so the disclosed total exceeds $500 million. Reporting at the time cited writedowns of as much as $569 million. Separately, the announcement was followed by a fall of roughly 45% in market value within a week, and the segment being closed had been 56% of revenue.
Was the Zillow algorithm inaccurate? No public figure exists for the accuracy of the purchase-decision model, which is the most useful number in the case and has never been published. The company's own explanation points at forecasting unpredictability, which is a statement about the market as much as the model. A valuation carrying a few percent of typical error is a perfectly good consumer product and a ruinous basis for a bid, so the error need not have been large to produce a nine-figure loss across a large inventory.
Why does the same model work as an estimate and fail as a purchase price? Because the acceptable error rate is a property of the use, not the model. An estimate given to a homeowner is information they combine with their own judgement, and the error is a caveat. The same number used to decide a purchase is a position, and the error becomes exposure. Moving an output from an advisory context to a transactional one changes the required accuracy by orders of magnitude without changing anything in the system.
What is selection bias in a decision model? Acting on a model's output selects for the errors that favour action. Buying when a model says a house is underpriced means buying disproportionately where the model was too high. Approving when a model says an application is low-risk means approving disproportionately where it underestimated. A model that is unbiased across all cases is not unbiased across the subset it causes you to act on, and this is nearly universal in decision systems and nearly never modelled.
Was this predictable? The structural difficulty was described in academic work published in December 2020, a year before the wind-down. Research on residential real estate intermediation set out the trade-off: attracting sellers who want liquidity requires paying near market value, which leaves a thin margin against the inventory risk taken on. That analysis was about the economics of the position rather than about machine learning, and it was legible to anyone reading the finance literature rather than the technology coverage.
What single control would have prevented it? A position limit. Capping how much inventory the model is permitted to acquire before someone reviews whether it is performing has nothing to do with machine learning and would have bounded the exposure regardless of how wrong the pricing turned out to be. Every other proposed fix depends on knowing in advance that the model was wrong, which is the thing nobody knew.
Did anyone outside the company get hurt? The losses fell on the company, its shareholders and roughly 2,000 employees affected by the workforce reduction. No homeowner was defrauded: sellers received the prices they had agreed to, and in many cases those prices turned out to be above what the market would later support. A securities class action was filed alleging inadequate disclosure during the period before the announcement.
Related articles
- Where AI has not landed: 77% report no use caseTransportation reports 7.5% AI use, construction 9.5%, against 73% for large information-sector firms. The most common reason given is not cost or skills. It is that no use case applies.
- AI in finance: the regulator looked and stepped backBanking has had formal model risk regulation since 2011. In April 2026 the successor framework arrived and deliberately placed generative AI outside its scope. That decision is the finding.
- AI in government: 126 use cases, 65 not made publicOne agency reported 126 active AI use cases and auditors found the inventory still incomplete, with tools contracted to build criminal cases missing from it entirely.
- AI incidents: two registers, 1,460 and 14,530The two main public AI incident databases count the same phenomenon and report 1,460 and 14,530. Neither is wrong. This is how to read an incident record, and what it cannot tell you.