Nine cases, and none was fixed by a better model
Territory 6 closes. Nine documented cases, three containing no AI at all, and not one where the remedy that worked was a more accurate system.
TL;DR. This record set out to document AI incidents against a stated standard: primary source, system identified, harm to someone other than the operator, causation stated at its actual strength. Nine cases in, three findings hold across all of them. Not one was resolved by a more accurate model; every remedy that worked was procedural or legal. Three of the nine contain no artificial intelligence at all, which means the failure mode does not require it. And the best-documented cases came from mechanisms built for something else entirely: a securities filing, a Royal Commission, a tribunal, a court. No AI-specific registry produced comparable evidence about any of them.
---
Status: synthesis. This article makes no new factual claims. Every figure appears in one of the nine case articles with its primary source, and each is linked at the point it is used.
---
The nine
| Case | Evidence | Harm | Remedy that worked |
|---|---|---|---|
| Moffatt v Air Canada | Tribunal decision | $650.88 | Contract law, unchanged |
| Amazon hiring | One news investigation | Disputed | Project cancelled |
| Zillow Offers | SEC filing | $304.4m, ~2,000 jobs | Business closed |
| Toeslagenaffaire | Parliamentary inquiry | 26,000 families | Government resigned |
| Williams v Detroit | Court record, settlement | 30 hours, three arrests | Lineup procedure banned |
| Epic sepsis model | Peer-reviewed validation | Uncountable | Local retraining advised |
| Horizon | High Court judgment, Act | Hundreds convicted | Convictions quashed by statute |
| Robodebt | Royal Commission | 470,000 debts | Scheme stopped, $1.8bn repaid |
| Tempe | NTSB investigation | A death | Testing suspended |
Finding one: the model was never the fix
Read the last column. Contract law, cancellation, closure, resignation, a procedural ban, a retraining recommendation, an Act of Parliament, a scheme shutdown, a suspension.
Not one of those is a more accurate system.
This is not because better models are impossible. It is because in every case the harm was determined by something downstream of the output.
In Moffatt the failure was two pages of the same website disagreeing. A consistency check, not a model, detects that.
In Zillow the estimate did not get worse; it moved from advising a homeowner to setting a purchase price, and the tolerance for its error changed by orders of magnitude with nothing in the model changing.
In Williams the face recognition system returned resembling faces, which is what it was built to do. The lineup assembled from its top candidate is what converted a similarity ranking into eyewitness identification. A model with half the error rate produces the identical rigged lineup on the cases it still gets wrong.
In the sepsis model the reported accuracy was not the operative number. On the only population an early-warning system exists for, cases clinicians had already missed, it contributed 7%.
Where the output enters a process built for a different kind of evidence, the process is the thing that has to change. Improving the output leaves the conversion intact.
Finding two: three of nine contain no AI
Horizon is transaction-processing software with defects. Robodebt is division. The Amazon case rests on a single news investigation and the operator disputes that anyone was affected at all.
Those three produce the same failure shape as the six that do involve models. An output trusted past its demonstrated reliability. A burden of proof that falls on the person least able to discharge it. Institutional certainty overriding contrary reports from people with no contact with each other.
If a mechanism reproduces itself in software containing no learned parameters, then the learned parameters are not the mechanism.
The practical consequence is uncomfortable for the field. Governance frameworks scoped to artificial intelligence will not catch these cases, because two of the three would fall outside the definition. A rule requiring model documentation, bias audits and explainability does nothing about a system that divides annual income by 26.
Finding three: the evidence came from elsewhere
The best-documented cases in this record were surfaced by mechanisms with no connection to AI oversight.
Securities disclosure produced the Zillow figure, because a material write-down must be reported to shareholders and misreporting is an offence.
A Royal Commission produced 990 pages on Robodebt, with subpoena power and sworn testimony.
Court records produced Moffatt, Horizon and Williams: a tribunal decision, a High Court judgment, a settlement with binding terms.
Peer review produced the sepsis validation, which nobody had commissioned and no regulation required.
And no AI-specific registry produced comparable evidence about any of them. The registers are useful for spotting patterns across many events and they are news-derived, which means they inherit what journalists noticed. As article 118 documented, roughly 15% of entries carry a harm classification and roughly 13% carry a cause.
Attaching an obligation to money, to liberty, or to sworn evidence produces better documentation than attaching it to technology. That observation should shape what anyone building AI accountability infrastructure actually builds.
Finding four: the burden of proof appears three times
Toeslagenaffaire, Horizon and Robodebt share a specific structure.
The institution asserts. The individual disproves. The evidence lives with the institution.
A Dutch family had to rebut a fraud determination to an agency with no discretion to reduce the consequence. A sub-postmaster had to show a system malfunctioned using logs held by the prosecutor. An Australian welfare recipient had to produce years of payslips against a departmental calculation.
Reversed proof multiplies every other flaw. A system with a modest error rate and a normal burden produces disputes that get resolved. The same system with the burden reversed produces payments, convictions and bankruptcies, because most people cannot discharge it and do not try.
This is the single highest-leverage design question in the record, and it is answerable before any model is trained: when this system is wrong about someone, who has to prove it?
Finding five: some harm cannot be counted
The sepsis model is the only case here that produces no artefact.
A patient whose sepsis is missed by both a clinician and an alert deteriorates, and the cause recorded is sepsis. Nothing in the file says a system failed to fire.
Every incident register counts events somebody noticed and reported. A warning system that quietly does not warn generates nothing to notice. That is not evidence the harm is small. It is evidence the instrument cannot see it, and it means the registers systematically over-represent failures that are visible and under-represent failures of omission.
Which is a serious problem, because omission is the failure mode of most assistive AI now being deployed.
What the record does not show
That AI is unusually dangerous. Nine cases is nine cases. There is no denominator here, no count of deployments that worked, and no basis for a rate.
That these are the worst cases. They are the best-documented ones, which is a different selection. Cases with equal harm and no court record, no filing and no inquiry are absent by construction.
That the pattern is causal. Nine cases selected for documentation quality, showing a common structure, is a hypothesis rather than a finding. The honest statement is that the mechanism recurs across every case where the documentation is good enough to see it.
And that anything here generalises to the frontier. Every case involves a deployed system doing a bounded task. None involves the capabilities that dominate current safety discussion.
What to do differently, in order
Decide what happens to the people the system is wrong about, before deciding how accurate it needs to be. The Dutch case turned on this: the same error rate under a proportionate recovery rule produces a dispute about a few hundred euros.
Put the burden on the institution. Where a system asserts something about a person, the institution should have to demonstrate the decision was sound, not the person that it was not.
Measure incremental contribution, not standalone accuracy. What the system adds beyond the process already running, on the cases that process misses.
Bound the exposure before you trust the output. A limit on how much a system may commit before someone reviews whether it works costs nothing and does not require knowing in advance that it is wrong.
And check whether the failure needs AI. If it does not, an AI governance programme will miss it.
The counter-argument
Nine cases cannot support five findings. This is a small, non-random sample selected for documentation quality, and the pattern may be an artefact of that selection. Cases resolved by better models would rarely produce a court record, so the claim that models are never the fix may reflect what generates paperwork rather than what fixes problems.
The synthesis understates the technology. Six of the nine do involve models, and in several the model's error was the necessary condition. Saying the fix was procedural is compatible with saying the cause was technical, and this article risks blurring the two.
Attributing the Horizon and Robodebt findings to AI governance stretches the category. If any automated system counts, this is a record about bureaucracy rather than about AI, and the boundary matters for whether the conclusions transfer.
And "measure incremental contribution" is easier to state than to do. It requires knowing what the unaided process would have caught, which usually requires a controlled comparison most organisations cannot run.
The short version
Nine documented cases, each with a primary source and a stated standard of inclusion.
Not one was resolved by a more accurate model. The remedies that worked were contract law, cancellation, closure, resignation, a procedural ban, a retraining recommendation, an Act of Parliament, a scheme shutdown and a suspension. Where an output enters a process built for a different kind of evidence, the process is what has to change.
Three of the nine contain no artificial intelligence. Horizon is software with bugs, Robodebt is division, and the Amazon case rests on a single investigation the operator disputes. They produce the same failure shape as the six that do involve models, which means the learned parameters are not the mechanism, and a governance programme scoped to AI will not catch them.
The best evidence came from mechanisms built for something else: securities disclosure, a Royal Commission, three courts, and one peer-reviewed validation nobody required. No AI-specific register produced comparable documentation of any case.
The burden of proof is reversed in three of them, and that is the single highest-leverage design question available: when this system is wrong about someone, who has to prove it? It is answerable before a model is trained.
And one case produces no artefact at all. A warning system that quietly does not warn generates nothing for a register to count, which means the record systematically under-represents failures of omission. That is the failure mode of most assistive AI now being deployed.
Common questions
What is the main finding of the incident record? That in nine documented cases, not one was resolved by a more accurate model. Every remedy that worked was procedural or legal: a contract law ruling, a project cancellation, a business closure, a government resignation, a ban on a lineup procedure, a retraining recommendation, an Act of Parliament, a scheme shutdown and a testing suspension. The harm in each case was determined by what happened to the output, not by how good the output was.
Why do three of the nine cases contain no AI? Because the failure mode does not require it. The Post Office Horizon system is conventional transaction-processing software with defects, Robodebt calculated debts by dividing annual income into equal fortnights, and the Amazon hiring case rests on a single news investigation whose central harm claim the operator disputes. All three produce the same shape as the cases involving models, which suggests the learned parameters are not the mechanism.
What does that imply for AI governance? That frameworks scoped to artificial intelligence will miss a substantial class of these failures, because the systems fall outside the definition. A rule requiring model documentation, bias auditing and explainability does nothing about a system that divides annual income by 26 and reverses the burden of proof. The scope that would catch all nine is automated decisions affecting people, not AI specifically.
Where did the best evidence come from? From mechanisms built for other purposes. Securities disclosure produced the Zillow write-down, because a material loss must be reported and misreporting is an offence. A Royal Commission produced 990 pages on Robodebt with subpoena power. Courts produced Moffatt, Horizon and Williams. Peer review produced the sepsis validation, which no regulation required. No AI-specific incident registry produced comparable documentation of any of these cases.
What is the single most useful design question? When this system is wrong about someone, who has to prove it? Three of the nine cases share a structure where the institution asserts, the individual disproves, and the evidence sits with the institution. Reversing the burden multiplies every other flaw, because most people cannot discharge it and pay, plead or comply instead. The question is answerable before any model is trained.
Why does the sepsis case matter differently? Because it produces no artefact. A patient whose sepsis is missed by both a clinician and an alert deteriorates, and the recorded cause is sepsis. Nothing says a system failed to fire. Incident registers count events somebody noticed, so a warning system that quietly does not warn is invisible to them. That is a limitation of the instrument rather than evidence the harm is small, and it means registers under-represent failures of omission, which is the failure mode of most assistive AI being deployed now.
Does this record show that AI is dangerous? No, and it cannot. Nine cases with no denominator supports no rate. There is no count here of deployments that worked, and the cases were selected for documentation quality rather than severity, so comparable harms without a court record, filing or inquiry are absent by construction. The record supports claims about mechanism, not about frequency.
What would change these conclusions? A well-documented case where the remedy that worked was a more accurate model would directly contradict the first finding. A systematic count of deployments with and without harm would allow rate claims the record currently cannot support. And an AI-specific oversight mechanism producing documentation comparable to a Royal Commission or a securities filing would undercut the third finding, which is at present an observation about where evidence actually comes from.
Sources
Primary documents only. Where a claim rests on a single report, the entry says so.
- The nine case articles in this record Artifipedia, Territory 6 This is a synthesis and makes no new factual claims. Every figure appears in one of the nine cases with its own primary source, and each is linked at the point it is used. The sources are theirs, not this article's.
Related articles
- The robots removed the walking. It was also the rest.The largest robot deployment on earth, and two sets of injury figures pointing in opposite directions. Both may be accurate, and what they share is more interesting than what they dispute.
- Horizon: the law presumed the computer was rightHundreds prosecuted on the output of an accounting system later found not to be robust. No machine learning was involved, which is precisely why it belongs in this record.
- The sepsis model caught 7% of what clinicians missedA sepsis warning system ran at hundreds of US hospitals before anyone outside the vendor validated it. The external check found the number that matters is not the one being reported.
- Robodebt: losing quietly to avoid losing publiclyA Royal Commission found the scheme unlawful, crude and cruel. The tribunal had been ruling against it for years, and the department never appealed, so no precedent was ever set.