Home/Blog/Safety & governance/Amazon's hiring AI: the case with no primary source
Amazon's hiring AI: the case with no primary sourceFour years from build to report. One source, five anonymous.201420162018build starts2014bias found2015project abandoned2017reported2018Four years from build to report. One source, five anonymous.
Four years from build to report. One source, five anonymous.

Amazon's hiring AI: the case with no primary source

The most-cited AI bias case in the world rests on one news investigation, five anonymous sources, no published numbers, and an operator that disputes the central harm claim.

TL;DR. A team began building a resume-scoring system in 2014. By 2015 it was rating candidates in a way that was not gender-neutral: it penalised the word "women's" and downgraded graduates of two all-women's colleges. The project was abandoned in 2017 and reported in October 2018. It is the canonical AI bias example, cited in textbooks, regulation debates and courtrooms. It also rests entirely on one news investigation with five anonymous sources. No technical report exists. No numbers were published. The two colleges are unspecified. The company states the tool was never used to evaluate candidates. By the standard this series set for the incident record, this is a near-miss with disputed causation, not a documented incident, and the gap between that and how it is cited is the most useful thing about it.

---

Status: reported, not established. Single-source: a Reuters investigation by Jeffrey Dastin, 10 October 2018, based on five people familiar with the effort, all anonymous. No primary document exists. The operator disputes the central claim about harm. Everything below is attributed accordingly.

---

The story as it is usually told: a team started work in 2014 on a system to score job applicants one to five stars, the way shoppers rate products. It was trained on ten years of resumes the company had received. Because the applicant pool skewed heavily male, the system learned that male candidates were preferable.

By 2015 the effect was visible. Resumes containing the word "women's", as in a women's chess club captaincy, were downgraded. So were graduates of two all-women's colleges. The system reportedly favoured verbs more common on male engineers' resumes.

The team edited the model to be neutral toward those specific terms. It could not guarantee the system would not find other proxies, and the project was abandoned in 2017.

That account is coherent, mechanistically plausible, and consistent with everything known about how these systems learn. It is also, in its entirety, one newspaper story.

What does not exist

This is the part almost never stated, and it matters more than the story.

No technical report. The company never published an analysis. There is no description of the architecture, the features, the training procedure or the evaluation.

No numbers. Not one. No selection rates by gender, no measured disparity, no candidate counts, no impact ratio. The most-cited example of algorithmic bias contains no measurement of bias.

No named institutions. The two all-women's colleges are unspecified in every version of the account. Nobody outside the company can check the claim.

No regulator involvement. No agency investigated, no finding was made, no enforcement action followed.

No identified affected person. Nobody has come forward as a candidate downgraded by this system, and given that it was reportedly never fully deployed, there may be no such person.

And the operator disputes the central point. The company's position is that the tool was never used by recruiters to evaluate candidates. A source described it as used only in a trial phase, never independently, never rolled out. Anonymous sources told Reuters recruiters did look at the ratings as one input among several. Those two accounts are not compatible, and no evidence exists that would settle them.

Applying the standard

Article 118 set six criteria for this record. This case is worth walking through them, because the result is not what its reputation implies.

Primary or near-primary source. Reuters is a serious outlet and Dastin is a serious reporter, and five sources is more than most investigations rest on. But an anonymous-source news story is not a court judgment, a regulator's finding or a published post-mortem. This is the weakest source category in the record.

System identified specifically enough to be checked. Partly. The operator is named and the purpose is clear. The system is not described in any way that could be verified.

Harm to someone other than the operator. Unestablished, and disputed. If the tool never evaluated candidates, the harm is to no one. If recruiters saw its ratings, some candidates were affected and none has been identified.

Causation stated at its actual strength. The reporting supports that the system produced gendered outputs. It does not establish that any hiring decision changed.

Near-misses labelled as such. On the available evidence, this is a near-miss, and it is almost universally cited as an incident.

The operator's account included. It is, above, and it directly contradicts the harm claim.

Verdict: this belongs in the record as a near-miss with disputed causation. That is a genuine entry and a much smaller claim than "Amazon's AI discriminated against women", which is how it appears in most citations.

Why it became canonical anyway

Four reasons, and none of them is evidentiary.

The mechanism is real and easy to explain. Train on historical outcomes from a skewed population and the model reproduces the skew. That is not in dispute, it is well documented elsewhere, and this case illustrates it memorably.

The detail is unforgettable. A system penalising the word "women's" is a perfect anecdote. It is concrete, it requires no technical background, and it survives retelling intact.

The company is enormous. A resume screener at a mid-sized firm would not have been reported. The name is doing much of the work.

And it arrived when the field needed an example. In 2018 the discussion about algorithmic bias was largely theoretical. This gave it a case.

A canonical example is chosen for being teachable, not for being well-evidenced, and this one is exceptionally teachable. That explains its position without justifying the weight placed on it.

What the case does establish

Being sceptical about the sourcing is not the same as dismissing it. Three things stand up.

Proxy discovery is real and hard to prevent. The reported behaviour, penalising a word that correlates with a protected characteristic rather than the characteristic itself, is exactly what these systems do. It is documented in the peer-reviewed literature independently of this case.

Removing one proxy does not remove the mechanism. The team reportedly neutralised the specific terms and abandoned the project because they could not be confident others were not present. That judgement is the correct one and it is the most valuable thing in the account. A model with access to text has access to thousands of unenumerated proxies.

And catching it before deployment is what success looks like. If the account is accurate, engineers found a problem in an internal tool and the company stopped. That is the outcome every governance framework is trying to produce, and it is routinely retold as a scandal.

Which is a perverse lesson to teach. An organisation that tests its own system, finds bias, fails to fix it and cancels the project has behaved well. Making it the standard cautionary tale creates an incentive not to look.

The one that should be cited instead

If the point is documented algorithmic hiring harm, better-evidenced cases exist.

Regulatory settlements. Enforcement actions against hiring-technology vendors produce named parties, findings and remedies. They are public and checkable.

Published bias audits. Under the New York City regime, employers must commission and publish independent audits with impact ratios. Those contain actual numbers, though as covered in the hiring article, only eighteen were found across 391 employers.

Peer-reviewed field experiments. Audit studies submitting matched applications through real screening systems produce measured disparities with confidence intervals.

None of those has the narrative force of a system that penalised the word "women's", which is why they are cited far less. The best-evidenced cases and the best-known cases are almost disjoint sets, and that is a fact about how examples spread rather than about the evidence.

The citation decay problem

There is a second failure here, downstream of the sourcing, and it is worth naming because it applies to every case in this record.

Retellings get stronger as they get further from the source.

The original reporting is careful. It attributes to people familiar with the effort, includes the company's denial, and states that the tool was used in trials. Each layer of retelling drops a qualifier.

The first layer drops the anonymity: Reuters reported that the system penalised women's resumes. The second drops the attribution: Amazon's AI penalised women's resumes. The third drops the trial status: Amazon used an AI that discriminated against women. The fourth supplies detail nobody has: Amazon's AI rejected qualified women for years.

By the fourth layer the claim is stronger than any evidence anyone has, and the chain of citations back to the source still looks intact. Every link cites a real thing that cited a real thing.

This is the mechanism by which a well-reported near-miss becomes a documented atrocity, and it operates without anyone lying at any step. It also operates fastest on cases that are most quotable, which selects precisely for the cases with the least underlying documentation, since documented cases come with numbers that constrain how far the retelling can drift.

The defence is mechanical: cite the primary source, not the citation. If the primary source is a news investigation with anonymous sources, say so, in the sentence where the claim appears rather than in a footnote.

That is why every entry in this record opens with a status line naming what it rests on. It is not academic decoration. It is the only thing that stops a claim getting stronger every time it is repeated.

What to take from it

Three practical things, none of which requires the story to be true in every detail.

Assume your model has proxies you have not enumerated. Any system with access to free text has access to thousands of correlates of protected characteristics. Removing the ones you thought of does not address the ones you did not.

Test before deployment, on your own pipeline. The reported failure was found in testing. Whatever else is uncertain, that part worked.

And be prepared to cancel. The hardest decision in the account is not detecting the bias, it is abandoning three years of work because the fix could not be verified. Most organisations do not do that, which is why most such systems ship.

What is unresolved

Whether any candidate was affected. The operator says no. Anonymous sources imply otherwise. There is no way to determine this from outside and there never will be.

What the actual disparity was. No figure has ever been published. The case is universally described as biased and nobody knows by how much.

Whether the project truly ended. Reporting mentions a reduced version retained for basic tasks. What that system does is not public.

And whether anything comparable is running now. Resume screening is widespread. This case is remembered because it was reported, and the reporting depended on five people choosing to talk.

The counter-argument

Anonymous-source reporting is how most corporate wrongdoing becomes known. Holding out for a primary document means waiting for organisations to publish their own failures, which almost never happens. Reuters verified with five independent sources, the company did not deny the technical account, and treating that as weak evidence sets a standard that would exclude most investigative journalism.

The operator's denial is narrower than it sounds. Saying a tool was never used to evaluate candidates is compatible with recruiters seeing its output. Carefully worded corporate statements are not the same as contradiction, and giving that denial equal weight to five sources may overcorrect.

The mechanism does not depend on the sourcing. Proxy discovery in text models is established independently. Even if every detail here were wrong, the lesson would survive, which is arguably why the case functions well as an example regardless of its evidentiary weight.

And near-miss versus incident may be the wrong distinction. A system that produced discriminatory scores inside a company is a real failure whether or not a candidate was rejected. Defining an incident by demonstrated individual harm excludes exactly the cases caught before harm, which are the ones worth studying.

The short version

A resume-scoring system was built from 2014, found by 2015 to be rating candidates in a way that was not gender-neutral, and abandoned in 2017. It penalised the word "women's" and downgraded graduates of two all-women's colleges. It is the canonical AI bias example.

It rests on one news investigation with five anonymous sources. No technical report. No published numbers of any kind, in the most-cited example of algorithmic bias. The two colleges are unspecified, so nobody outside the company can check. No regulator investigated. No affected person has been identified. And the operator states the tool was never used to evaluate candidates, which anonymous sources contradict and no evidence can settle.

Against the criteria this record uses, that makes it a near-miss with disputed causation. A real entry, and a much smaller claim than the one usually made from it.

Three things stand up regardless. Proxy discovery is real and documented independently. Removing one proxy does not remove the mechanism, and the team's judgement that they could not verify the absence of others was correct. And catching it before deployment is what success looks like.

That last point is the uncomfortable one. An organisation that tested its own system, found bias, could not fix it and cancelled the project behaved well. Retelling that as the standard cautionary tale creates an incentive not to look.

Common questions

What happened with Amazon's AI recruiting tool? According to a Reuters investigation published in October 2018, a team began building a resume-scoring system in 2014 that rated candidates one to five stars. Trained on ten years of resumes from a heavily male applicant pool, by 2015 it was rating candidates in a way that was not gender-neutral: it penalised resumes containing the word "women's" and downgraded graduates of two all-women's colleges. The team neutralised those specific terms, could not be confident other proxies were absent, and the project was abandoned in 2017.

Is the Amazon hiring bias case well documented? No, and this is rarely stated. The entire account rests on one news investigation citing five anonymous sources. There is no technical report, no published figures of any kind, no named institutions, no regulatory finding and no identified affected candidate. Reuters is a serious outlet and five sources is substantial for an investigation, and it remains the weakest source category: not a court judgment, a regulator's finding or a published post-mortem.

Did Amazon's AI actually reject women? Unestablished. The company states the tool was never used by recruiters to evaluate candidates, and a source described it as used only in a trial phase, never independently and never rolled out. Anonymous sources told Reuters that recruiters did look at the ratings as one input among several. Those accounts are incompatible and no evidence exists that would settle them. No individual has been identified as affected.

Why is this case cited so much if the evidence is thin? Because it is teachable rather than because it is well-evidenced. The mechanism is real and easy to explain, the detail about penalising the word "women's" is unforgettable and needs no technical background, the company is enormous enough to make it newsworthy, and it arrived in 2018 when discussion of algorithmic bias was largely theoretical and needed a case. The best-evidenced cases and the best-known cases are almost disjoint sets.

What is proxy discrimination? A model finding features that correlate with a protected characteristic without using the characteristic itself. Reported examples here include a word associated with women's activities and the names of women's colleges. It is well documented in the peer-reviewed literature independently of this case, which is why the mechanism stands regardless of the sourcing. Any system with access to free text has access to thousands of unenumerated correlates.

Why couldn't they just fix the bias? They reportedly did fix the specific terms and could not establish that others were absent. That is the correct judgement and the most valuable part of the account. Removing an enumerated proxy does not remove the mechanism that finds proxies, and with free-text input the space of possible correlates is not something anyone can fully search.

What should companies learn from it? Assume your model has proxies you have not enumerated, since removing the ones you thought of does not address the ones you did not. Test before deployment on your own pipeline, which is the part of this account that worked. And be prepared to cancel: the hardest decision reported here was abandoning three years of work because a fix could not be verified, and most organisations do not make it, which is why most such systems ship.

Are there better-documented hiring discrimination cases? Yes. Regulatory settlements against hiring-technology vendors produce named parties, findings and remedies. Bias audits published under the New York City regime contain actual impact ratios, though only eighteen were found across 391 employers surveyed. And peer-reviewed audit studies submitting matched applications through real systems produce measured disparities with confidence intervals. None has the narrative force of a system that penalised the word "women's", which is why they are cited far less.

Learn the concepts

← All posts