Williams v Detroit: the match was not the failure
The first publicly reported wrongful arrest from a face recognition match. The settlement's remedy is procedural, and it names a failure that has nothing to do with model accuracy.
TL;DR. In January 2020 Detroit police arrested Robert Williams outside his home, in front of his wife and two young daughters, for a 2018 shop theft he had nothing to do with. The lead came from a face recognition search run on a blurry still from surveillance video. He was held thirty hours. It was the first publicly reported case of a false face recognition match producing a wrongful arrest. The lawsuit settled in June 2024 for $300,000 and, more consequentially, a set of binding policy changes. The settlement does not require a more accurate system. It prohibits arresting on a match alone, and it prohibits building a photo lineup out of a face recognition result. That second prohibition names the actual failure, and it is not a model problem.
---
Status: established. Primary sources: the case record in Williams v. City of Detroit and the settlement agreement of 28 June 2024, both documented by the ACLU and the University of Michigan Civil Rights Litigation Initiative, who brought the case. Figures below are from that record.
---
In 2018 someone took several watches from a Shinola store in Detroit. Investigators pulled a blurry, low-quality still from the store's surveillance video and sent it to Michigan State Police to run through face recognition.
The search returned Robert Williams' driver's licence photograph.
In January 2020 he was arrested at his home in Farmington Hills, in front of his family and his neighbours, and detained for thirty hours in an overcrowded cell. He was not the man in the video and had not been near the store.
The lawsuit, filed in 2021, alleged Fourth Amendment and Elliott-Larsen Civil Rights Act violations, and that the detective's warrant application omitted enough to mislead the magistrate into finding probable cause that was not there.
Discovery established two things the city had not disclosed. Detroit had no policy governing law enforcement use of face recognition at the time. And it had not trained officers on how the technology fails.
What the settlement actually requires
This is the part worth reading closely, because it is not what most coverage implies.
No arrest on a face recognition result alone. A match is a lead. It is not identification and cannot carry an arrest by itself.
And no photo lineup built directly from a face recognition search. A witness may not be shown an array assembled from the system's own candidates.
The second is the one that matters, and it is the mechanism behind every one of Detroit's wrongful arrests.
Here is why. A face recognition system returns the faces in the database that most resemble the probe image. Put that top candidate into a six-person lineup with five fillers, and you have not tested the witness. You have shown them a picture the machine already selected for resemblance, surrounded by people it did not select. The witness picks the one that looks most like their memory, which is exactly the one the algorithm chose for looking most like the image.
The lineup then launders an algorithmic guess into eyewitness identification, which courts and juries treat as strong evidence. The identification is not independent. It never was.
Nothing about a more accurate model fixes this. A system with half the error rate produces the same rigged lineup on the cases it still gets wrong.
The measurement came first
A detail that changes how this case should be read.
In December 2019 the US National Institute of Standards and Technology published its evaluation of demographic effects in face recognition, part of the Face Recognition Vendor Test programme. Across a large number of algorithms it found false positive rates varying substantially by demographic group, with higher rates for several groups including Black and Asian faces, and the pattern holding across most of the algorithms tested.
Williams was arrested the following month.
The differential error rate was not discovered afterwards, and it was not disclosed by an investigation into the arrest. It had been measured and published by the federal government's own testing body, before the arrest, using the vendors' own systems.
All three of Detroit's publicly reported wrongful arrests from this technology were of Black residents.
Four things this establishes
A lead is not an identification, and systems that return ranked candidates invite the confusion. The output is "these faces resemble the probe." It is read as "this is the person." The gap between those is where the harm sits.
Procedural remedies can bind where technical ones cannot. The settlement does not specify an accuracy threshold, because a threshold would need auditing, would change with every model update, and would still permit exactly the failure that occurred. Prohibiting a use is enforceable in a way that requiring a performance level is not.
Absence of policy is a finding, not a gap. Detroit had deployed the technology with no rules and no training. That was not an oversight discovered later; it was the operating condition, and it only surfaced through litigation discovery.
And a downstream procedure can destroy the independence of evidence. The lineup was the failure. The system merely supplied its input.
What it does not establish
That face recognition cannot be used. The settlement permits it as an investigative lead and constrains what may be done with that lead. Williams and the ACLU both said they would prefer it not be used at all, and both said the settlement represents the strongest constraints achieved.
That the specific match was caused by demographic error. A blurry still from surveillance video will produce false matches regardless of who is in it. The NIST findings establish that error rates differ by group; they do not establish which factor produced this particular match.
That other departments are bound. A settlement binds its parties. Detroit's policy is the strongest of its kind in the United States and it applies to one police department.
And that the count is complete. These are the publicly reported cases. A wrongful arrest that ends without charges, or without a lawyer who identifies the technology in the file, does not enter any register.
The pattern across this record
Four cases in, something consistent is showing up.
Moffatt turned on a publication contradicting itself, not on a model error. Zillow turned on an estimate moving from advice to a bid without its error bar changing. The Dutch benefits scandal turned on a recovery rule with no proportionality, applied to selections a model made. And this one turns on a lineup procedure that stripped the independence out of an eyewitness identification.
In none of the four is the fix a better model. In all four the harm is set by what the institution does with the output.
That is not a coincidence of case selection. It is what happens when a system's output enters a process built for a different kind of evidence, and nobody changes the process.
What is unresolved
How many cases there are. Detroit's became public because Williams found counsel who recognised what had happened. Cases where the technology is not named in the file are not counted anywhere.
Whether the policy holds. The settlement includes audit provisions, and whether they are exercised over years is not yet known.
Whether other departments follow. No mechanism requires them to. Adoption elsewhere has been voluntary and uneven.
And what the current error rates are. NIST's evaluation is ongoing and algorithms have improved since 2019. Whether the differential has narrowed, and by how much, is a live question with published answers that change.
The counter-argument
The technology worked as specified. It was asked which database faces resemble a probe image and it returned them, ranked. It did not identify anyone. Every subsequent step, the lineup, the warrant application, the arrest, was performed by people who treated a similarity ranking as an identification. Calling this a face recognition failure locates the error in the only component that did its job.
A blurry frame is the harder problem. The probe was a low-quality still from store surveillance. Any identification method operating on that input would have a high error rate, and the fault may lie more in accepting the probe than in the matching.
The settlement may be too narrow. It constrains procedure without limiting deployment, and a department that follows the letter while treating a match as effectively conclusive can still reach the same outcome by a slightly longer route.
And it may be too broad. Barring lineups built from face recognition results removes an investigative tool in cases where it would produce correct identifications. The cost of that is not zero and is not measured anywhere, which cuts both ways in an article about unmeasured costs.
The short version
Detroit police arrested Robert Williams in January 2020 for a 2018 theft, on a face recognition lead generated from a blurry surveillance still. He was held thirty hours, in front of his family and neighbours. It was the first publicly reported wrongful arrest from a false face recognition match. The case settled in June 2024 for $300,000 and binding policy changes. Discovery established Detroit had no policy for the technology and no officer training.
The settlement does not require a more accurate system. It prohibits arrest on a match alone, and it prohibits building a photo lineup from a face recognition search.
That second prohibition names the failure. A face recognition system returns faces that resemble the probe. Put its top candidate in a lineup with five fillers and the witness is not being tested: they are shown the face a machine already selected for resemblance, surrounded by faces it did not. The lineup launders an algorithmic guess into eyewitness identification, which courts treat as strong evidence. A model with half the error rate produces the identical rigged lineup on the cases it still gets wrong.
And the differential error rate was public first. NIST published its demographic evaluation of face recognition in December 2019, finding false positive rates that varied substantially by group across most algorithms tested. Williams was arrested the month after. All three of Detroit's publicly reported wrongful arrests were of Black residents.
Four cases into this record the pattern is consistent. A publication contradicting itself, an estimate that became a bid, a recovery rule with no proportionality, and a lineup that destroyed the independence of a witness. In none of them is the fix a better model.
Common questions
What happened to Robert Williams? Detroit police arrested him at his home in January 2020, in front of his wife and two young daughters, for a 2018 theft of watches from a Shinola store. The lead came from a face recognition search run on a blurry, low-quality still taken from the store's surveillance video, which returned his driver's licence photograph. He was detained for thirty hours. He was not the man in the footage and had not been near the store. It was the first publicly reported instance of a false face recognition match leading to a wrongful arrest.
What did the settlement require? It was announced on 28 June 2024 and included a $300,000 payment and binding policy changes. Detroit police may not arrest anyone based solely on a face recognition result, and may not conduct a photo lineup assembled directly from a face recognition search. The agreement is described by the parties who brought it as the strongest constraints on police use of the technology adopted by any US department.
Why does the lineup prohibition matter more than the accuracy question? Because it names the mechanism. A face recognition system returns the database faces that most resemble a probe image. Placing its top candidate in a lineup with five fillers does not test the witness: it presents a face already selected by a machine for resemblance, alongside faces that were not. The witness picks the closest match to their memory, which is the one the algorithm chose. That converts an algorithmic guess into eyewitness identification, which courts treat as strong evidence. A more accurate model produces the same rigged lineup on the cases it still gets wrong.
Was the technology known to be less accurate for some groups? Yes, and before the arrest. The US National Institute of Standards and Technology published an evaluation of demographic effects in face recognition in December 2019, as part of its Face Recognition Vendor Test programme, finding false positive rates that varied substantially between demographic groups across most of the algorithms tested. Williams was arrested the following month. All three of Detroit's publicly reported wrongful arrests from the technology were of Black residents.
Did Detroit have rules for using it? No. Discovery in the lawsuit established that the department had no policy governing law enforcement use of face recognition at the time of the arrest, and had not trained officers on how the technology fails. That absence was not disclosed beforehand and surfaced only through litigation.
Does this mean police should not use face recognition? The settlement does not say that. It permits use as an investigative lead and constrains what may be done with the lead. Williams and the ACLU have both said they would prefer it not be used at all and that this demand did not succeed, and that the resulting constraints are nonetheless the strongest in the country. That distinction between what was sought and what was achieved is worth preserving when the case is cited.
How many wrongful arrests like this have there been? Nobody knows. Three have been publicly reported in Detroit alone, all of Black residents, and others have been reported elsewhere. A wrongful arrest that ends without charges, or where no one identifies the technology in the case file, does not enter any register. The publicly reported cases are a floor, not a count.
What is the transferable lesson? That a downstream procedure can destroy the independence of evidence, and no amount of model improvement addresses it. The system returned a ranked list of resembling faces, which is what it was built to do. Every step after that treated a similarity ranking as an identification. Where a model's output enters a process designed for a different kind of evidence, the process is the thing that has to change.
Sources
Primary documents only. Where a claim rests on a single report, the entry says so.
- Williams v. City of Detroit American Civil Liberties Union, case record and 2024 settlement The case record, the discovery findings on absent policy and training, and the settlement terms.
- Civil Rights Advocates Achieve the Nation's Strongest Police Department Policy on Facial Recognition Technology ACLU of Michigan, 28 June 2024 The settlement announcement, including the prohibition on lineups assembled from a face recognition search.
- Face Recognition Vendor Test Part 3: Demographic Effects, NISTIR 8280 US National Institute of Standards and Technology, December 2019 The federal evaluation finding false positive rates that varied substantially by demographic group across most algorithms tested. Published the month before the arrest. Named by report number rather than linked, because NIST publication URLs have moved.
Related articles
- The Dutch benefits scandal: the rule, not the modelAround 26,000 families were wrongly accused of fraud and a government resigned. The parliamentary inquiry did not blame the algorithm. What it found is more useful and less quoted.
- Horizon: the law presumed the computer was rightHundreds prosecuted on the output of an accounting system later found not to be robust. No machine learning was involved, which is precisely why it belongs in this record.
- Robodebt: losing quietly to avoid losing publiclyA Royal Commission found the scheme unlawful, crude and cruel. The tribunal had been ruling against it for years, and the department never appealed, so no precedent was ever set.
- AI in hiring: 18 bias audits from 391 employersResearchers checked 391 New York employers against the world's first algorithmic bias audit law. Eighteen had posted an audit. Nearly every audit that existed reported passing.