COMPAS: both sides of the dispute were correct
A newspaper said a risk score was biased. The vendor said it was fair. Two research teams then proved independently that both claims were true and cannot both be fixed.
TL;DR. In May 2016 ProPublica reported that a recidivism risk score used in Broward County misclassified Black defendants as high risk at 1.9 times the rate of white defendants among people who were not rearrested. The vendor responded that the tool satisfied predictive parity: a given score meant the same probability of reoffending regardless of race. Both analyses were correct. Within a year, two teams proved independently that when base rates differ between groups, calibration and equal error rates cannot both hold. The algebra does not bend. Which means fairness is not a property a model can have. It is a choice between incompatible definitions, and the choice is normative rather than technical.
---
Status: established. Primary sources: the ProPublica investigation of 23 May 2016 and the dataset it published, the vendor's published rebuttal, and two peer-reviewed impossibility results, Kleinberg, Mullainathan and Raghavan (2016) and Chouldechova (2017). The dispute is a matter of public record and so is its resolution.
---
COMPAS is a risk assessment instrument. It produces a score intended to indicate the likelihood that a defendant will be rearrested, and courts have used those scores to inform bail, parole and supervision decisions.
In May 2016 ProPublica published an analysis of scores for defendants in Broward County, Florida. Among people who were not rearrested in the following two years, Black defendants were 1.9 times more likely than white defendants to have been labelled high risk. Among people who were rearrested, Black defendants were substantially less likely to have been labelled low risk.
Those are false positives and false negatives, and they were unequal by race.
The vendor responded with a different measurement. Its analysis showed the instrument satisfied predictive parity: among defendants given a particular score, the proportion who went on to be rearrested was approximately the same regardless of race. A score of seven meant the same thing whoever received it.
Neither party was misrepresenting the data. They were measuring different quantities and each found what they measured.
The result that settled it
Within roughly a year, two teams working separately established why the dispute could not be resolved by better analysis.
Kleinberg, Mullainathan and Raghavan showed that three natural fairness conditions cannot be satisfied simultaneously except in degenerate cases. Chouldechova proved the specific version at issue: calibration and error-rate balance cannot coexist when the two groups have different base rates.
The mechanism is arithmetic rather than statistical. If one group is rearrested at a higher rate than another, and a score is calibrated so that it means the same thing for both, then applying any single threshold to that score produces different false positive and false negative rates between the groups. You may choose which quantity to equalise. You cannot equalise both.
So the dispute was not about who had analysed the data correctly. Both had. It was about which definition of fairness to adopt, and that question has no empirical answer.
Working the arithmetic, because the result sounds like a trick
The impossibility is easy to state and easy to disbelieve, so it is worth walking through with numbers. The figures below are illustrative, chosen for clean arithmetic rather than taken from the case.
Two groups of 1,000 people each. Group A has a base rate of 30%: 300 will be rearrested. Group B has a base rate of 50%: 500 will be. Suppose the instrument is perfectly calibrated, so anyone scored high risk has a 60% chance of rearrest in either group.
To be calibrated at 60%, the high-risk group must contain 60% true cases and 40% false ones, in both groups.
Group A has 300 true cases to draw from. Suppose 200 of them are scored high risk. Calibration then requires about 133 false positives alongside them, giving 333 high-risk labels of which 200 are correct. Those 133 false positives come from the 700 people who were not rearrested, a false positive rate of about 19%.
Group B has 500 true cases. Suppose 400 are scored high risk. Calibration requires about 267 false positives, giving 667 labels of which 400 are correct. Those 267 come from the 500 who were not rearrested, a false positive rate of about 53%.
Same score, same meaning, same threshold. 19% against 53%.
The second group's false positive rate is higher because it has fewer true negatives to spread the same proportion of errors across. Nothing about the model produced that. The base rates did.
Now try to fix it. Lower the threshold for Group A and you break calibration: a high-risk label now means something different depending on group. Raise it for Group B and the same thing happens in reverse. Every adjustment that equalises one measure moves the other.
Which is why this is not a bug anyone can be blamed for and not a problem better engineering solves. It is a property of applying one threshold to two populations with different prevalence, and it would hold for a hand-written rule, a human assessor applying consistent standards, or a perfect oracle. The only escapes are to abandon a single threshold, to abandon calibration, or to change what is being predicted.
Why this is the most important case in the record
The other entries describe failures: a wrong answer, a bad rule, a model used outside its tolerance. This one describes something harder, which is a system working correctly under one reasonable definition and unacceptably under another, with no configuration that satisfies both.
Three consequences follow.
A claim that a model "is fair" is incomplete. It means the model satisfies some criterion, and unless the criterion is named the claim carries no information. Anyone asserting fairness without specifying which definition has either chosen one silently or not checked.
Fairness audits measure a choice, not a property. An audit reporting that a system passes has reported that it passes the test that audit selected. A different auditor with a different criterion could examine the same system and fail it, both correctly.
And the choice is a value judgement made by whoever picks the metric. Equalising false positives means fewer people wrongly detained from the higher-base-rate group, at the cost of the score meaning different things by group. Equalising calibration means the score is consistent, at the cost of unequal wrongful classification. Both are defensible positions about what a criminal justice system should prioritise, and neither is a technical finding.
What the impossibility rests on, and why it matters
The proof requires unequal base rates. That condition deserves examination rather than acceptance.
The base rate here is not offending. It is rearrest. The outcome variable in these datasets is whether a person was arrested again within a period, which is a function of behaviour and of policing: where officers patrol, which offences are pursued, and who is stopped.
So the mathematics is downstream of a measurement decision. The impossibility is real given the data. What it establishes is that no fairness criterion can be jointly satisfied on a label that itself carries the pattern of the enforcement that produced it.
That is a different and more uncomfortable finding than "you must trade off fairness definitions." It says the trade-off is forced by a quantity nobody chose to measure and everyone treats as ground truth. Improving the model cannot address it. Only changing what is predicted, or what the prediction is used for, can.
What ProPublica did that deserves more credit
One detail is consistently underweighted.
ProPublica published the dataset. The scores, the defendant records and the outcomes were obtained through public records requests, matched across sources, and released.
That is the reason the dispute could be resolved at all. The vendor could run its own analysis. Academics could test both claims. The impossibility results could be demonstrated against the actual data rather than a hypothetical. Subsequent work could revisit the dataset and identify problems with it, which is itself only possible because the data existed publicly.
Almost no comparable investigation does this. A finding published without its data produces an argument. A finding published with its data produces a field, and algorithmic fairness as a research area substantially dates from this exchange.
It is worth separating that from whether the headline conclusion was right. The methodology was reproducible, which is the higher standard and the rarer one.
What the case does not establish
That COMPAS was inaccurate. Its overall predictive performance was comparable across groups. The dispute concerned the distribution of errors, not the quantity.
That the tool was designed to discriminate. Race was not an input. The disparity arises from correlated features and unequal base rates, which is what makes the finding structural rather than a matter of intent.
That risk assessment should not be used. The comparison is not against perfection but against unaided judicial discretion, which is unaudited, unmeasured and varies between individuals. Whether a measurable instrument with known disparities is better or worse than an unmeasurable process with unknown ones is a real question and this case does not answer it.
And that the impossibility means fairness is hopeless. It means formal parity criteria conflict. It does not mean nothing can be improved, and a body of later work argues the formal framing itself is the limitation.
How to use this when evaluating a system
Five things, and the first is the one that would have prevented the entire dispute.
Name the criterion before you measure. Decide which definition of fairness the system is accountable to, in writing, before evaluating it. Choosing afterwards means choosing the one it passes.
Report the others anyway. A system calibrated by design will have unequal error rates when base rates differ. Publishing both numbers is honest and costs nothing, and it prevents the next investigation being a revelation.
Ask what the label actually measures. Rearrest is not offending. Default is not inability to pay. Attrition is not performance. Every impossibility argument is conditioned on a base rate, and the base rate belongs to a measurement someone chose.
Ask who bears each error. The trade-off is not abstract. Equalising one metric moves harm from one group to another, and the question of which harm matters more is for the people accountable for the system rather than the people tuning it.
And ask what the score is used for. A risk score informing a supervision level and the same score informing detention carry the same errors at entirely different cost, which is the finding from the benefits case arriving in a different jurisdiction.
What is unresolved
Whether formal criteria are the right frame at all. A substantial line of work argues that satisfying parity metrics is not the same as producing just outcomes, and that the impossibility results define the problem too narrowly. That debate is live.
What the data would show with a better label. No large-scale risk instrument has been validated against actual offending rather than rearrest, because that measurement does not exist.
Whether disclosure changed practice. Risk assessment remains widely used. Whether the debate altered how instruments are validated, or mainly produced a literature, is not clearly established.
And what the counterfactual is. No study has compared outcomes under algorithmic risk assessment against outcomes under the discretionary process it partly replaced, at scale, over time. That is the comparison that matters and it has not been made.
The counter-argument
Error rate imbalance may be the less relevant metric here. Several researchers argued at the time that the disparity ProPublica measured is an expected consequence of differing prevalence and does not by itself indicate a biased instrument. On that reading the investigation identified a real statistical property and attached the wrong interpretation to it.
The vendor's position was the mainstream statistical one. Predictive parity is what a well-calibrated instrument is supposed to deliver, and criticising a tool for achieving its design goal is an odd basis for a finding of bias. That the finding was widely reported as proof of a racist algorithm outran what the analysis supported.
The impossibility results can be over-read. They establish that specific formal criteria conflict under specific conditions. They are frequently cited to imply that fairness is unachievable in general, which is a considerably stronger claim than anything proved.
And the practical question was never the metric. What matters is whether a defendant is detained who should not have been. That depends on the threshold, the discretion available to the judge, and what detention does to a person, none of which the fairness debate addressed. The argument about which parity criterion applies consumed a decade of attention that the question of what the scores were used for did not receive.
The short version
In May 2016 ProPublica reported that among Broward County defendants who were not rearrested within two years, Black defendants had been labelled high risk at 1.9 times the rate of white defendants, with the mirror disparity among those who were rearrested. The vendor responded that the instrument satisfied predictive parity: a given score carried the same probability of rearrest regardless of race.
Both analyses were correct. Within a year Kleinberg, Mullainathan and Raghavan, and separately Chouldechova, proved that calibration and error-rate balance cannot both hold when base rates differ between groups. You may choose which to equalise. Not both.
Which makes this the most consequential case in the record, because it is not a failure. It is a system working correctly under one reasonable definition and unacceptably under another, with no configuration satisfying both.
Three things follow. A claim that a model is fair carries no information unless the criterion is named. A fairness audit measures a choice rather than a property, and a different auditor could correctly reach the opposite verdict. And the choice is a value judgement made by whoever selects the metric, not a technical result.
And the condition the impossibility rests on deserves more scrutiny than it gets. The base rate is not offending, it is rearrest, which is a function of behaviour and of policing. The trade-off is forced by a quantity nobody chose to measure and everyone treats as ground truth. No model improvement addresses that.
One thing deserves more credit than it receives. ProPublica published the dataset, which is why the vendor could respond, academics could test both claims, and the impossibility could be demonstrated against real data. A finding published without its data produces an argument. A finding published with its data produced a field.
Common questions
What was the COMPAS controversy? In May 2016 ProPublica analysed recidivism risk scores for defendants in Broward County, Florida, and found that among people not rearrested within two years, Black defendants had been labelled high risk at 1.9 times the rate of white defendants. The vendor responded that the instrument satisfied predictive parity, meaning a given score carried the same probability of rearrest regardless of race. Both analyses were correct measurements of different quantities.
Was COMPAS biased or not? It depends entirely on which definition of fairness is applied, and that is not a question data can settle. It failed error-rate balance, meaning false positive and false negative rates differed by race. It satisfied predictive parity, meaning a score meant the same thing for everyone who received it. Two research teams proved independently in 2016 and 2017 that when base rates differ between groups, no instrument can satisfy both.
What is the fairness impossibility theorem? The result that several natural fairness criteria cannot be satisfied simultaneously. Kleinberg, Mullainathan and Raghavan showed three conditions conflict except in degenerate cases. Chouldechova proved that calibration and error-rate balance cannot coexist when groups have different base rates. The mechanism is arithmetic: if a score means the same thing for both groups and one group has a higher base rate, any single threshold produces different error rates between them.
What is the difference between calibration and error-rate balance? Calibration, or predictive parity, asks whether a given score carries the same probability of the outcome for every group. Error-rate balance asks whether the false positive and false negative rates are the same for every group. Calibration looks at the population from the score outward; error-rate balance looks from the outcome back to the score. Both are reasonable, and when base rates differ they are mathematically incompatible.
Does this mean fairness is impossible? No, and this is the most common over-reading. The results establish that specific formal parity criteria conflict under specific conditions. They do not establish that nothing can be improved, and a substantial line of later work argues that formal parity is the wrong frame entirely and that just outcomes require examining the system a tool sits inside rather than the tool's error distribution.
Why does the base rate matter so much? Because the impossibility only holds when base rates differ. Here the base rate is rearrest within a period, which is a function of behaviour and of policing patterns: where officers patrol, which offences are pursued, and who is stopped. The trade-off is therefore forced by a measurement decision nobody deliberately made, and improving the model cannot address it. Only changing what is predicted, or what the prediction is used for, can.
What should a company take from this case? Name the fairness criterion your system is accountable to in writing before you evaluate it, since choosing afterwards means choosing the one it passes. Report the other criteria anyway, because a calibrated system will have unequal error rates when base rates differ and publishing both is honest and free. Ask what your label actually measures. Ask who bears each kind of error. And ask what the score is used for, since identical errors carry entirely different cost depending on the consequence attached.
Why is ProPublica's data release significant? Because it made the dispute resolvable. The scores and outcomes were obtained through public records requests, matched across sources, and published. That allowed the vendor to run its own analysis, academics to test both claims, the impossibility results to be demonstrated against real data, and later researchers to identify problems with the dataset itself. Almost no comparable investigation publishes its data, and algorithmic fairness as a research field substantially dates from this exchange.
Related articles
- Amazon's hiring AI: the case with no primary sourceThe most-cited AI bias case in the world rests on one news investigation, five anonymous sources, no published numbers, and an operator that disputes the central harm claim.
- AI in government: 126 use cases, 65 not made publicOne agency reported 126 active AI use cases and auditors found the inventory still incomplete, with tools contracted to build criminal cases missing from it entirely.
- What a model card should say and usually does notA model card was meant to be an evidentiary document: what a model was trained on, where it fails, how it was measured. Regulation has now made most of it mandatory, and the gap between the template and its honest use is the whole subject.
- The Dutch benefits scandal: the rule, not the modelAround 26,000 families were wrongly accused of fraud and a government resigned. The parliamentary inquiry did not blame the algorithm. What it found is more useful and less quoted.