AI in law: 1,313 filings sanctioned in 106 countries
A researcher has catalogued 1,313 court proceedings involving AI-fabricated content, 496 involving licensed attorneys. Single-matter sanctions went from $5,000 to $55,597 in two years.
TL;DR. Law is the only profession with a public, adversarial, quantified record of AI failure, because filings are public and opponents have a professional incentive to check them. A database maintained by a researcher at HEC Paris had catalogued 1,313 court proceedings involving AI-fabricated content by April 2026, across 106 countries, 496 involving licensed attorneys. The count was around 200 a year earlier. Sanctions have escalated from $5,000 in the first case to $55,597 in a single matter. What works is retrieval over a closed corpus you control. What fails is generation of citations, and the failure mode has evolved from obviously invented cases to fabricated quotations attributed to real ones, which are far harder to catch.
---
In June 2023 a New York lawyer filed a brief containing six case citations that did not exist. He had asked a chatbot for supporting authority and filed what it produced. The court imposed $5,000 in sanctions and called the circumstance unprecedented.
It was not unprecedented for long.
By April 2026 a database maintained by a research fellow at HEC Paris had catalogued 1,313 court proceedings in which AI-generated fabrications were submitted to courts, across 106 countries, with 496 involving licensed attorneys rather than self-represented litigants. The count stood at roughly 200 a year earlier and around 719 in January. Documented cases are being added at five to six per day.
Single-matter sanctions have reached $55,597, an eleven-fold escalation in eighteen months. US courts imposed at least $145,000 in the first quarter of 2026 alone.
No other profession has a record like this, and the reason is structural rather than incidental. Legal filings are public, and an opposing counsel has both the skill and the motive to check every citation. Law is not worse at using AI than medicine or finance. It is the only field where the failures are systematically discovered, documented and quantified by an adversary.
Why this is the best-evidenced domain in the series
Worth stating plainly, because it changes how to read the numbers.
In most fields, an AI error is absorbed. A wrong marketing claim is a wasted campaign. A wrong clinical suggestion is caught by a clinician, or it is not caught and never attributed. A wrong code suggestion fails a test or ships as a bug nobody traces back.
In litigation, an error is examined by someone paid to find it. Opposing counsel reads the brief. A clerk pulls the cited authority. When the citation does not resolve, the discovery is entered into the public record, with the attorney's name attached, in a document that will exist permanently and be searchable.
That adversarial structure converts private failures into public data. The 1,313 cases are not evidence that law has an AI problem others do not. They are evidence that law has a detection system others do not, and the honest reading is that comparable error rates exist elsewhere and go unmeasured.
There is a second consequence. Because courts publish reasoning, the profession has accumulated doctrine faster than any other. A hospital learns from an incident internally. A court writes an opinion that other courts cite.
What actually works
Three categories, and the distinction between them is one property.
Retrieval over a corpus you control
The largest genuine use, and it predates the current wave by two decades.
Electronic discovery is the search of large document sets for material relevant to a matter. Predictive coding, where a model trained on attorney-reviewed samples ranks the remaining documents, has been accepted by courts since the early 2010s and is routine.
It works because the corpus is fixed, closed and yours. Every document the system surfaces exists, because it came from the production set. The model is ranking, not generating. A false positive costs a wasted review; a false negative is a known and quantified risk that sampling protocols are designed to bound.
Drafting where the source is supplied
Summarising a deposition, drafting a first-pass contract clause from a template library, or producing a chronology from documents already in evidence.
Same property: the material exists before the model touches it. The model reorganises rather than invents, and the reviewing attorney is checking a transformation of something they can compare against.
Search over a licensed legal database
Where a vendor grounds generation in an actual case-law corpus with citation links, the failure mode narrows considerably. It does not vanish. One assessment found that around 24% of frontier-model legal answers cite law that does not support the proposition they are offered for, which is a different and subtler failure than inventing a case outright.
The common property across all three: the authority exists independently of the model, and the lawyer can verify it against a source the model did not produce. Everything that has gone wrong has gone wrong where that property was absent.
The failure record, and how it evolved
The interesting part is not that fabrication happened. It is what happened next.
Phase one was invented cases. A citation to a case that does not exist, with a plausible name, a plausible reporter, a plausible year. Embarrassing, and comparatively easy to catch: the citation does not resolve. This is the Mata pattern, and it is what most people picture.
Phase two is harder. As one federal appellate court has described it, hallucinations now frequently take the form of fabricated quotations from real cases, or mischaracterised holdings. The case exists. The reporter citation resolves. The pinpoint page is real. The quoted sentence is not in it, or the case held something adjacent to what it is cited for.
That is a materially worse problem, and the reason is procedural. A verification step that checks whether a citation resolves will pass a fabricated quotation from a real case. The lawyer confirms the case is real, sees the citation is correctly formatted, and files. Catching the second phase requires reading the cited authority, which is the thing the tool was supposed to save time on.
So the failure mode adapted to defeat the most common verification practice. Not deliberately, but the effect is the same, and any verification protocol written for phase one is now inadequate.
What the courts have settled
Doctrine has accumulated quickly, and four points are now well established across jurisdictions.
Using AI is not itself sanctionable. Filing unverified output is. The duty is candour toward the tribunal and competence in representation, and neither is suspended because a tool was involved. Courts have been consistent that the obligation attaches to the signature on the filing.
Candour after the fact substantially affects the penalty. One federal appellate court stated directly that had the attorney accepted responsibility and been more forthcoming, lesser sanctions would likely have followed. Another characterised an attorney's explanations as attenuated and treated the failure to be forthcoming as strongly favouring sanction. The compounding error is the response, not the filing.
Supervision extends to the tool. The rules governing responsibility for the work of non-attorney assistants have been applied to AI output. A partner who did not personally file is not thereby insulated.
And the conclusions are converging internationally. Courts in London, Singapore, Vancouver and multiple Argentine provincial appellate courts have reached comparable positions through independent reasoning between 2023 and 2026. That convergence across legal traditions suggests the doctrine is tracking something structural about the technology rather than reflecting one jurisdiction's preferences.
Professional guidance has followed. The American Bar Association issued its first formal ethics opinion on generative AI in July 2024, covering competence, confidentiality, candour, supervision and fees, after earlier state-level guidance.
The asymmetry nobody has resolved
The most uncomfortable finding, and the one that gets least attention.
A survey by researchers at Northwestern found that over 60% of federal judges report using AI tools themselves.
Those are the same courts imposing sanctions on attorneys for insufficient verification. The standard being enforced on filings is not obviously the standard being applied to the bench, and no framework currently exists that addresses it.
This is not an accusation. Judicial use is mostly for summarisation and drafting rather than for finding authority, which is the lower-risk category identified above. But the asymmetry is real, unexamined, and would be a serious problem if a judicial opinion were found to contain a fabricated citation. A profession enforcing a verification standard it has not formally applied to itself is in an unstable position, and the instability has not been addressed by any bar or judicial conference.
The detection asymmetry, generalised
The finding that law is not worse but merely watched has an implication for every other domain in this series, and it is worth extracting before moving on.
Error rates are only knowable where someone is paid to find errors.
Litigation has an adversary. Every filing is read by a party whose interest is served by finding a mistake in it, and who has the training to spot a citation that does not resolve. So the error rate becomes a public number.
Medicine has an adversary only after harm. Malpractice litigation surfaces errors that produced injury, years later, filtered by whether anyone connected the outcome to the decision. A diagnostic suggestion that was wrong and harmless is never recorded anywhere.
Software has an adversary that is not a person. Tests and compilers catch a specific class of error mechanically and comprehensively, and catch nothing outside it. A function that passes its tests and implements the wrong requirement ships.
Most domains have no adversary at all. A marketing claim, a summarised report, a translated document, a customer service answer. Nobody checks, so nobody knows, and the absence of documented failures reads as an absence of failures.
Which means the 1,313 number should be read as a floor on what a well-instrumented domain finds, not as a ceiling on what AI gets wrong. Law looks bad in this data because law is looking. The fields that look clean are mostly the fields that are not.
The practical form: when evaluating AI in any domain, ask who checks the output, what they are checking for, and what class of error their check cannot see. If the answer is that nobody checks systematically, an absence of reported failures tells you nothing at all.
Adoption, and the gap inside it
The usage numbers explain why the failure count keeps climbing.
Reported adoption has risen sharply. One professional survey found lawyers using AI-based tools rising from 11% in 2023 to 30%, with a sharp size divide: 46% at firms of 100 or more attorneys against 18% among solo practitioners. Another found active generative AI use among legal organisations moving from 14% to 26% in a single year, with 78% of law firm respondents expecting it to be central within five years.
And the governance gap: 52% of professionals reported their organisation still had no policy covering it.
Put those together and the trajectory is legible. Adoption is rising faster than governance, the sanctioned cases are concentrated among practitioners without institutional verification processes, and the size divide suggests the difference is resource rather than judgement. A large firm has a librarian, a citation-checking workflow and a partner who will not sign an unverified brief. A solo practitioner has a deadline.
How to use it without being sanctioned
Six rules, drawn from what the sanctioned cases have in common.
Never let a model supply an authority. Ask it to explain a case you found. Ask it to summarise a document you have. Do not ask it what supports your argument, because that is the request that produces fiction.
Verify the quotation, not the citation. Phase-two failures pass a citation check. Open the case and find the sentence. If the tool saved you less time than that costs, it was not saving you time.
Check the holding, not just the existence. A real case cited for a proposition it does not support is a misrepresentation to the tribunal whether or not a model was involved.
Keep the verification record. The cases that resolved best involved attorneys who could show what they checked and when. The ones that resolved worst involved explanations offered after the fact.
If you find an error, say so immediately and completely. This is the single highest-value rule in the list, because courts have said explicitly that candour reduces sanctions and that attenuated explanations increase them.
And write the policy before you need it. Fifty-two percent of organisations have not. The sanctioned cases cluster where no process existed, which means the process is the intervention rather than the tool choice.
What is unresolved
Whether the phase-two failure rate is measurable at all. Fabricated quotations from real cases are caught only when someone reads the authority. The 1,313 documented cases are, by construction, the ones that were found. Nobody knows the denominator, and the phase-two shift means the detection rate is probably falling as the failure gets subtler.
Whether grounded legal research tools solve it. Vendors offering retrieval over licensed case-law corpora report much lower fabrication rates, and independent assessment has still found around a quarter of frontier-model legal answers citing law that does not support the proposition. Whether that number falls with better retrieval or reflects something harder is not established.
What happens to junior lawyers. First-pass research and document review are how junior practitioners learn to read authority critically. If those tasks are automated, the skill that catches a fabricated quotation may not develop in the people who will need it most.
And whether the bench asymmetry gets addressed. Over 60% judicial usage against an enforced verification standard for attorneys is not a stable arrangement, and no judicial conference has published a framework for it.
The counter-argument
The count is a denominator problem and reads worse than it is. 1,313 proceedings sounds enormous and is a vanishingly small fraction of filings worldwide over three years. Against millions of documents filed, a four-figure error count may represent a lower rate than pre-AI citation errors, which were never systematically tracked because nobody thought to look.
The comparison class is not perfection. Lawyers cited cases incorrectly before generative AI. Misquotation, mischaracterised holdings and citations to overruled authority are old problems with a long professional literature. What is new is the volume and the traceability, not the category of error.
Sanctions are a sign the system is functioning. A profession that detects a new failure mode, publishes reasoning about it, converges internationally within three years and escalates penalties proportionately is responding well. Reading the sanction count as evidence of crisis inverts what it actually demonstrates.
And the working uses are large and boring. Predictive coding in discovery has been court-accepted for over a decade and processes volumes no team of associates could read. An article organised around fabricated citations risks implying the technology has no place in law, when the accurate statement is that one specific use, asking a model for authority, is the one that fails.
The short version
A New York lawyer filed six non-existent case citations in June 2023 and was sanctioned $5,000. By April 2026 a database maintained at HEC Paris had catalogued 1,313 court proceedings involving AI-fabricated content across 106 countries, 496 involving licensed attorneys. The count was around 200 a year earlier and around 719 in January, growing at five to six documented cases a day. Single-matter sanctions have reached $55,597, an eleven-fold escalation in eighteen months, with at least $145,000 imposed by US courts in the first quarter of 2026.
Law is not worse at using AI than other professions. It is the only one where the failures are found. Filings are public and opposing counsel has both the skill and the motive to check every citation, which converts private error into public data. The honest reading is that comparable error rates exist in fields with no adversary checking.
Three uses work, and they share one property: retrieval over a closed corpus you control, drafting where the source material is supplied, and search grounded in a licensed case-law database. In each, the authority exists independently of the model. Everything that has gone wrong has gone wrong where it did not.
The failure mode has evolved and this is the part most verification protocols have not caught up with. Phase one was invented cases, which fail a citation check. Phase two is fabricated quotations from real cases and mischaracterised holdings, which pass one. Catching those requires reading the cited authority, which is precisely the work the tool was supposed to remove.
Courts have settled four points quickly: using AI is not sanctionable but filing unverified output is, candour after the fact materially reduces the penalty, supervision duties extend to the tool, and the conclusions have converged across London, Singapore, Vancouver and Argentine appellate courts through independent reasoning.
And the asymmetry nobody has addressed: over 60% of federal judges report using AI tools themselves, while their courts enforce a verification standard on attorneys that no framework has formally applied to the bench.
Common questions
How many lawyers have been sanctioned for AI-generated citations? A database maintained by a research fellow at HEC Paris had catalogued 1,313 court proceedings involving AI-fabricated content by April 2026, across 106 countries, of which 496 involved licensed attorneys rather than self-represented litigants. The count stood at roughly 200 a year earlier and around 719 in January 2026, with five to six new documented cases being added per day.
What was the first AI sanctions case? A New York matter decided in June 2023, in which an attorney filed a brief containing six fabricated case citations produced by a chatbot. The court imposed $5,000 in sanctions and described the circumstance as unprecedented. It became the reference case cited by subsequent courts, including a Massachusetts decision in February 2024 that quoted it directly.
How large are AI citation sanctions now? Single-matter sanctions have reached $55,597, an escalation of roughly eleven times in eighteen months from the $5,000 imposed in the first case. US courts imposed at least $145,000 in the first quarter of 2026 alone, including a record penalty in Oregon and the first substantial federal appellate fine linked to AI-tainted briefs.
Is it against the rules for a lawyer to use AI? No. Courts have been consistent that using AI is not itself sanctionable and that filing unverified output is. The duties engaged are candour toward the tribunal and competence in representation, neither of which is suspended because a tool was involved. The American Bar Association issued its first formal ethics opinion on generative AI in July 2024, covering competence, confidentiality, candour, supervision and fees.
What does AI actually do well in legal work? Three things, sharing one property. Retrieval over a closed document set you control, as in electronic discovery, where predictive coding has been court-accepted for over a decade. Drafting where the source material is supplied, such as summarising a deposition already in evidence. And search grounded in a licensed case-law corpus. In each the authority exists independently of the model, which is exactly what fails when a model is asked to supply authority itself.
Why do AI legal hallucinations still happen if everyone knows about them? Because the failure mode changed. Early fabrications were wholly invented cases, which fail a check of whether the citation resolves. Current fabrications frequently take the form of invented quotations attributed to real cases, or mischaracterised holdings, which pass that check entirely. Catching them requires reading the cited authority, which is the work the tool was supposed to save.
How do I verify AI output for a court filing? Never let the model supply an authority; ask it to explain a case you found rather than to find one. Verify the quotation rather than the citation, since a fabricated quote from a real case passes a citation check. Confirm the holding supports the proposition. Keep a record of what you verified and when. And if you find an error, disclose it immediately and completely, because courts have said explicitly that candour reduces sanctions and attenuated explanations increase them.
Are judges using AI too? Yes. A survey by researchers at Northwestern found over 60% of federal judges report using AI tools, mostly for summarisation and drafting rather than for locating authority. This creates an unresolved asymmetry, since the same courts enforce verification standards on attorneys that no judicial conference has formally applied to the bench, and it would become a serious problem if a published opinion were found to contain a fabricated citation.
Related articles
- AI in government: 126 use cases, 65 not made publicOne agency reported 126 active AI use cases and auditors found the inventory still incomplete, with tools contracted to build criminal cases missing from it entirely.
- AI in hiring: 18 bias audits from 391 employersResearchers checked 391 New York employers against the world's first algorithmic bias audit law. Eighteen had posted an audit. Nearly every audit that existed reported passing.
- AI in finance: the regulator looked and stepped backBanking has had formal model risk regulation since 2011. In April 2026 the successor framework arrived and deliberately placed generative AI outside its scope. That decision is the finding.
- AI in medicine: 1,524 devices, 1.6% with trial dataThe FDA lists 1,524 AI-enabled medical devices. A review of 691 found 1.6% cited a randomised trial and under 1% reported patient outcomes. Medicare pays for about ten.