Home/Blog/Evaluation & evidence/AI incidents: two registers, 1,460 and 14,530
AI incidents: two registers, 1,460 and 14,530Same phenomenon. The gap is definition, not error.14,530monitor, incidents + hazards1,460curated databaseSame phenomenon. The gap is definition, not error.
Same phenomenon. The gap is definition, not error.

AI incidents: two registers, 1,460 and 14,530

The two main public AI incident databases count the same phenomenon and report 1,460 and 14,530. Neither is wrong. This is how to read an incident record, and what it cannot tell you.

TL;DR. Two public registers track AI failures. One records 1,460 incidents, the other 9,218 incidents plus 5,312 hazards. The gap is not error, it is definition: one counts alleged harm including near-misses, the other separates events that caused harm from events that could have. Almost every entry in both originates in a news article, so the record measures what was reported, which is not what happened. Only about 15% of entries in the larger-curated database carry any structured harm classification at all. And the monitor with the higher count uses language models to classify AI incidents, with no published error rate for the classifier. This piece sets the method for everything that follows in this series.

---

There are two widely used public registers of AI failures, and they disagree by an order of magnitude.

One records 1,460 incidents. The other records 9,218 incidents alongside 5,312 hazards, a combined 14,530.

Neither is wrong. They are counting different things and using the same word for both.

The first defines an incident as an alleged harm or near-harm event to people, property or the environment where an AI system is implicated. Alleged is doing real work in that sentence: the entry can stand on a credible report without the causal claim being established.

The second defines an incident as an event where the development, use or malfunction of AI systems directly or indirectly leads to harm, and defines a hazard separately as an event that could plausibly lead to one. Splitting the two is more precise and it means any count depends on which filter is applied.

So before any figure from either can be used, you need to know which definition produced it, whether near-misses are inside or outside, and whether the number includes hazards. A citation of "over 14,000 AI incidents" that omits the second half is not a citation, it is a rounding of two categories into one.

Where the entries come from, which bounds everything

The single most important structural fact about both registers, and the one least often stated.

Almost every entry originates in a news article. One is built from community submissions and curated news reports, the other from clusters of articles supplied by a news intelligence platform. Public submission is accepted by both and is a minority of volume.

That has a specific consequence. The record measures what was reported. What was reported measures what was newsworthy. And newsworthiness is a function of novelty, identifiable victims, a nameable company and a journalist's beat.

Four biases follow, and they are structural rather than fixable.

Language and geography. The top source domains are large English-language outlets. A failure in a jurisdiction those outlets do not cover is not in the record.

Nameable defendants. An incident involving a recognisable company is reported. The same failure in a system nobody has heard of is not.

Discrete over diffuse. A wrongful arrest has a person, a date and a name. A recommendation system slightly degrading a million decisions has none of those and appears nowhere.

And novelty decay. The first fabricated legal citation was news. The thousandth is not, so the record's growth rate for any failure type falls as the failure becomes normal, independent of whether it is becoming more or less common.

Which means an incident count is a measure of attention, and its rate of change is a measure of attention changing. That is truly useful and it is not what most people quoting these numbers believe they are quoting.

What is inside the record, and how little is typed

The registers publish taxonomies for classifying harm type and failure cause. The coverage is thinner than the existence of the taxonomies suggests.

In the curated database, roughly 214 entries carry a harm classification and 188 carry a failure-cause classification, against a total of about 1,460. That is around 15% and 13%.

So the great majority of the record is a title, a date, a set of source links and free text. It supports counting and it does not support most of the analysis people attempt on it.

This is not a criticism of the maintainers. Classifying an incident against a taxonomy requires reading the sources, forming a judgement about causation from incomplete third-party reporting, and recording how confident that judgement is. One taxonomy addresses this directly with confidence modifiers, which is the right design and it is slow.

The practical consequence: any claim of the form "X% of AI incidents are caused by Y" is being computed over the classified minority unless stated otherwise, and the classified minority was selected by whoever had time to classify it.

The classifier problem

The higher-count monitor identifies entries by retrieving AI-tagged events from a news platform and then using language models to classify them as incidents, hazards or unrelated. A smaller model filters, a larger one confirms.

That is a reasonable engineering decision at that volume, and it introduces something worth naming. The public register of AI failures is populated by an AI classifier whose own error rate is not published.

Both directions matter. False positives inflate the count with events that are not incidents. False negatives remove events that are, invisibly, because nothing records what the filter rejected.

Neither is knowable from outside, and it means the difference between 1,460 and 14,530 is partly a definitional gap and partly the difference between human curation and automated classification at scale. Nobody has published a decomposition.

What the composition figures say

With the caveats above attached, the reported distribution is informative about attention if not about frequency.

Growth runs at roughly 35 to 45% year on year, faster than AI deployment growth. That could mean failures are outpacing adoption, or that reporting is catching up, or that the category of things counted as AI has widened. All three are consistent with the number.

Generative AI accounts for around 58% of recent entries, which tracks both deployment and newsworthiness.

Severity is stable, with about 3% classified as fatal or major harm. The stability across a period of rapid growth is the more interesting part: whatever is driving the count is not driving severity.

By type, misinformation and content harms run near 28%, discrimination and bias near 22%, physical safety near 14%. The first category expanded substantially after 2023, which is a real change in what is being reported and not necessarily in what is occurring.

The gap nobody has closed

Analyses of AI incident reporting keep returning the same four gaps, and they are worth stating plainly because they explain why the registers look as they do.

No standard definition. The two main registers use different ones. Regulatory instruments use others. Whether a near-miss qualifies and whether actual harm is required are unsettled across jurisdictions.

No standard reporting format. Nothing specifies what fields an incident report contains, so entries are not comparable across sources.

No assessment procedure. There is no agreed method for deciding whether an AI system caused an outcome, which is the question every entry turns on and the hardest one to answer from news coverage.

And no incentive to disclose. This is the binding constraint. An organisation that reports its own AI failure receives regulatory attention, reputational damage and possible liability. One that does not receives nothing. Every register is therefore built almost entirely from failures that became public against the operator's interest.

One study of production incidents in generative AI cloud services found 38.3% were reported by humans rather than caught by automated monitoring. That is inside organisations with proper observability. The fraction reaching a public register from anywhere is far smaller.

How to read any incident count

Six questions, and they apply to any figure quoted from any register.

Which definition? Alleged harm, established harm, or harm plus hazards. The answer changes the number by roughly a factor of ten.

Does it include near-misses? Both registers handle these differently and the choice is rarely stated when the number is quoted.

How was it classified? Human curation and automated classification produce different populations from the same underlying events.

What fraction is typed? If the claim is about causes or categories, it is computed over the classified subset, which is a minority.

What would not appear? Diffuse harm, unnamed operators, non-English jurisdictions and failure modes that have stopped being novel.

And what is the denominator? An incident count without a deployment count says nothing about rate. Twelve hundred incidents against ten thousand deployments and against ten million are different worlds, and nobody knows the second number.

Why this series maintains a record anyway

Given all of the above, the case for another one has to be specific.

Not to be comprehensive. Two registers with funding and staff cannot achieve that, and a third would not.

To write up the small number of incidents that are fully documented, properly. Most entries in the large registers are a headline and a link. A court judgment, a regulator's finding or a published post-mortem supports something better: what the system was, what it did, what the consequence was, what changed afterwards, and what the record does not establish.

To state what each case does and does not show. The most common misuse of an incident is as evidence for a general claim it cannot support. A single sanctioned filing is not a fabrication rate.

And because the pattern across a well-documented set is the useful output. Twelve well-sourced cases with their mechanisms stated support an argument that fourteen thousand headlines do not.

Every entry in this series will name its sources, state what is established and what is alleged, and say plainly what it fails to demonstrate. Where a case rests on a single report, that will be said. Where the operator disputes the account, that will be said too.

The inclusion standard this series uses

Stating it up front, because a record whose criteria are unstated is an opinion with citations attached.

An entry requires a primary or near-primary source. A court judgment, a regulator's finding, a published post-mortem, a filed complaint, a company statement, or reporting by an outlet that names its documents. A single story citing an unnamed source is not enough on its own and will be marked as such if included.

The system must be identified specifically enough to be checked. "An AI system" is not an entry. What it was, who operated it, and what it was deployed to do are the minimum, and where a vendor is disputed rather than confirmed that is stated.

The harm must be to someone other than the operator. A company losing money on its own bad model is a business outcome. The record is for consequences that landed on people who did not choose the system.

Causation is stated at its actual strength. Established, alleged, disputed, or unknown. Most public incidents sit in the second or third category, and writing them as the first is the most common failure in this genre.

Near-misses are included and labelled. They are informative about mechanism and they are not evidence of harm, so they are counted separately rather than folded into a total.

And the operator's account is included where one exists. A disputed incident with both positions stated is more useful than a clean one with only the accusation.

What this excludes is as important as what it admits. Not included: capability demonstrations with no deployment, red-team findings without a production system, model outputs a researcher elicited deliberately, and anything where the only source is a screenshot. Each of those is interesting and none of them is an incident.

And one honest limitation. Applying these criteria means the record will be small. Most publicly discussed AI failures do not have a primary source behind them. A register of a dozen properly documented cases is more useful than a thousand headlines, and it is also a much less impressive number, which is the trade being made deliberately.

What is unresolved

Whether mandatory reporting will change the picture. Several jurisdictions are drafting incident reporting obligations. If they arrive, the registers become a different kind of object, and the transition will look like an explosion in incidents that is entirely an artifact of disclosure.

How to count diffuse harm. No proposed framework handles a system that makes a million decisions slightly worse. It is plausibly the largest category of real harm and it is structurally invisible to every register.

Whether the classifier gap can be measured. Publishing precision and recall for the automated classification would let anyone decompose the difference between the two counts. Nobody has.

And whether incident counts should be used for policy at all. They are currently cited in regulatory debate as evidence of trend. Given that they measure reported attention, using them to set thresholds risks regulating the news cycle.

The counter-argument

Imperfect records are how every safety field started. Aviation incident reporting began with inconsistent voluntary accounts and became the most effective safety instrument in any industry. Criticising early AI registers for lacking standardisation describes their age rather than their value, and the alternative is nothing.

The definitional gap is a feature. Two registers with different thresholds serve different users. A regulator wanting established harm and a researcher wanting near-misses need different filters, and collapsing them into one standard would serve one poorly.

News-sourcing has a real virtue. It is adversarial. A journalist verifying a story applies scrutiny that self-reported incident data does not receive, which is the same argument that made legal filings the best-documented domain in the previous series.

And automated classification is the only way to cover the volume. Human curation produced 1,460 entries in several years. If the true number is larger, an imperfect classifier that finds most of them is more useful than perfect curation that finds a tenth, and the honest response is to publish its error rate rather than abandon the method.

The short version

Two public registers track AI failures. One records 1,460 incidents. The other records 9,218 incidents and 5,312 hazards. Neither is wrong: the first counts alleged harm including near-misses, the second separates harm that occurred from harm that could have. Any figure quoted from either is uninterpretable without knowing which definition produced it.

Almost every entry originates in a news article, so the record measures what was reported, which measures what was newsworthy. Four biases follow structurally: English-language coverage, nameable operators, discrete over diffuse harm, and novelty decay that makes any failure type appear to slow as it becomes normal. An incident count is a measure of attention.

Roughly 15% of entries in the curated register carry a harm classification and 13% a failure-cause classification. Any claim about what causes AI incidents is computed over that minority.

And the register with the higher count uses language models to classify AI incidents, with no published error rate. The difference between the two totals is part definitional and part the difference between human curation and automated classification, and nobody has published a decomposition.

The four gaps identified repeatedly in the literature are no standard definition, no standard format, no assessment procedure, and no incentive to disclose. The last is binding: an organisation reporting its own failure receives scrutiny and liability, one that stays quiet receives nothing, so every register is built from failures that became public against the operator's interest.

This series maintains a record for a narrow reason. Not to be comprehensive, which two funded registers cannot manage. To write up the small number of cases that are properly documented, state what each does and does not establish, and let the pattern across a well-sourced set carry the argument that fourteen thousand headlines cannot.

Common questions

How many AI incidents have there been? Nobody knows, and the two main public registers differ by roughly a factor of ten. One records 1,460 incidents. The other records 9,218 incidents plus 5,312 hazards, a combined 14,530. The gap is definitional rather than an error: the first counts alleged harm including near-misses, the second separates events that caused harm from events that could plausibly have caused one.

What counts as an AI incident? There is no standard answer. One widely used definition is an alleged harm or near-harm event to people, property or the environment where an AI system is implicated. Another is an event where the development, use or malfunction of AI systems directly or indirectly leads to harm, with hazards defined separately as events that could plausibly lead to an incident. Whether near-misses qualify and whether actual harm is required remain unsettled across jurisdictions.

Where does incident data come from? Almost entirely from news articles. One register curates community submissions and reported cases; the other retrieves AI-tagged event clusters from a news intelligence platform. This bounds what can appear: coverage skews to English-language outlets, to incidents with a nameable operator, and to discrete harms with identifiable victims. Diffuse harm across many small decisions is structurally invisible.

Are AI incidents increasing? Reported incidents grow at roughly 35 to 45% year on year, faster than AI deployment growth. That figure is consistent with three different explanations: failures outpacing adoption, reporting catching up with existing failures, or the category of things counted as AI widening. The data does not distinguish them, and severity has stayed roughly stable at around 3% classified as fatal or major harm.

What kinds of AI incidents are most common? By reported share, misinformation and content harms run near 28%, discrimination and bias near 22%, and physical safety near 14%, with generative AI accounting for around 58% of recent entries. The first category expanded substantially after 2023, which reflects a change in what is being reported and not necessarily in what is occurring.

Why don't companies report their own AI failures? Because nothing rewards it. An organisation that discloses receives regulatory attention, reputational damage and possible liability, while one that stays quiet receives nothing. This absent incentive is identified repeatedly as the binding constraint on incident reporting, and it means public registers are built almost entirely from failures that became public against the operator's interest.

Is AI used to classify AI incidents? Yes, in the register with the higher count. AI-tagged events are retrieved from a news platform, a smaller language model filters out unrelated items, and a larger one confirms whether each remaining event is an incident or a hazard. It is a reasonable choice at that volume and the classifier's error rate is not published, so neither false positives inflating the count nor false negatives silently removing entries can be assessed from outside.

How should I cite an AI incident count? State which register, which definition, and whether hazards are included, since those three choices move the number by about a factor of ten. If the claim concerns causes or categories, note that classification covers a minority of entries. And avoid using a count as a rate: an incident total without a deployment total says nothing about how often systems fail, and the deployment denominator is unknown.

Learn the concepts

← All posts