Home/Blog/Safety & governance/AI in hiring: 18 bias audits from 391 employers
AI in hiring: 18 bias audits from 391 employers391 employers checked. 18 posted an audit. Nearly all of those passed.100%employers checked5%posted an auditOF ORGANISATIONS DEPLOYING AI391 employers checked. 18 posted an audit. Nearly all of those passed.
391 employers checked. 18 posted an audit. Nearly all of those passed.

AI in hiring: 18 bias audits from 391 employers

Researchers checked 391 New York employers against the world's first algorithmic bias audit law. Eighteen had posted an audit. Nearly every audit that existed reported passing.

TL;DR. New York City passed the world's first law requiring independent bias audits of automated hiring tools, effective July 2023. Researchers then checked 391 employers: eighteen had posted an audit report and thirteen had posted the required transparency notice. Nearly every audit that did exist reported passing. Neither number means what it appears to, because the law lets employers decide whether their own tool is in scope, so a missing audit cannot be distinguished from a tool that was declared out of scope. The researchers named this null compliance, and it is the most important concept in algorithmic regulation, because every framework built since has the same structure.

---

In July 2023 New York City became the first jurisdiction anywhere to require independent bias audits of commercial algorithmic systems. Local Law 144 obliges any employer using an automated employment decision tool to have it audited annually for race and sex bias, publish the results, and notify candidates ten business days in advance.

Researchers at Cornell and Data & Society then went and checked. 155 investigators recorded compliance across 391 employers.

Eighteen had posted an audit report. Thirteen had posted a transparency notice.

That is 4.6% and 3.3%.

And the finding that makes those numbers uninterpretable rather than merely bad: nearly every audit that was published reported an impact ratio above 0.8, the threshold conventionally treated as passing. So the visible picture is that almost nobody audits, and almost everybody who audits passes.

Why the low number is not simply non-compliance

The researchers were careful about this, and the care is the contribution.

The law grants employers substantial discretion over whether a given tool falls within its scope. An automated employment decision tool is defined as a computational process that issues a score, classification or recommendation substantially assisting or replacing a discretionary employment decision.

Substantially assisting is doing the work in that sentence, and the employer decides.

A resume-parsing system that surfaces candidates for human review can be characterised as informing a decision rather than substantially assisting it. A ranking that a recruiter is free to ignore can be described the same way. Neither characterisation is obviously wrong, and both remove the tool from scope.

So an absent audit has two indistinguishable explanations: the employer is not complying, or the employer determined in good faith that the law does not apply. The researchers call this null compliance, and it means the observed 4.6% cannot be read as an enforcement failure or as evidence that most employers do not use these tools. The measurement does not separate them.

That is not a flaw in the study. It is a flaw in the law, faithfully measured.

Why the high pass rate is not reassuring

The second finding compounds the first.

Nearly all published audits reported an impact ratio above 0.8. That figure comes from the four-fifths rule, a long-standing rule of thumb in United States employment discrimination practice: if the selection rate for one group falls below four-fifths of the rate for the most-selected group, that is treated as evidence of adverse impact worth investigating.

An audit reporting above 0.8 is reporting no adverse impact by that measure.

Two readings are available and they lead to different places.

The optimistic reading: hiring tools are mostly fine. Vendors have known about disparate impact for years, they test for it, and the audits confirm the testing worked.

The structural reading: employers who expect to fail do not publish. If you can define your tool out of scope, and you would rather not publish a failing number, the rational move is to determine that the law does not apply. The published set is not a sample of tools. It is a sample of tools whose owners chose to publish.

Nothing in the data distinguishes these. But the second is what the incentive structure predicts, and a regulation that produces a 96% publication gap alongside a near-100% pass rate among publishers has produced exactly the pattern selection bias produces.

The law does not require fixing anything

The point most commonly misunderstood, and it is worth being exact.

Local Law 144 does not prohibit a biased tool. It requires an audit, publication of the result, and notice to candidates. A tool with a poor impact ratio is not thereby illegal under this law.

It may create exposure elsewhere. Title VII and the New York City Human Rights Law both address discriminatory employment practice, and a published audit showing adverse impact is evidence a plaintiff would very much like to have. But the audit regime itself is a transparency mechanism, not a remediation one.

This is a deliberate design choice and it has a consequence: the law's theory of change is that publication creates pressure. That theory requires someone to read the publications, which brings us to the third finding.

The researchers also assessed the value of the regime to actual job seekers and found it limited, because of shortcomings in accessibility and usability. Audit reports are posted on employer websites in formats and locations that a candidate is unlikely to find, in a statistical vocabulary they are unlikely to parse, describing a tool they may not know was used.

A transparency mechanism that the intended beneficiary cannot practically use is a disclosure regime rather than an accountability one.

What an impact ratio does and does not tell you

Worth understanding, because it is the number every one of these audits turns on.

The calculation is straightforward. Take the selection rate for each demographic group, meaning the fraction who passed the tool's screen. Divide each by the rate for the highest-scoring group. The result is the impact ratio, and it is computed across sex, race and ethnicity, and their intersections.

It measures outcomes, not mechanism. A tool can achieve an acceptable ratio while using a proxy that correlates with a protected characteristic, provided the net effect happens to balance. It can also fail the ratio while using no problematic feature at all, if the applicant pool itself is skewed.

It is a snapshot of the data it was run on. An audit conducted on one quarter's applicants for one role at one company says something about that. It does not transfer to a different role, a different pool or a different quarter, and the annual cadence means a tool can drift for eleven months before anyone measures again.

And the four-fifths rule is a rule of thumb, not a legal standard. It originated as an enforcement screening device, and courts have treated it as one input rather than a test. An audit reporting 0.81 has not proved anything; it has failed to trip a heuristic.

The regulator noticed

In early 2026 the New York City Comptroller published a critical audit of the city's own enforcement, reporting major gaps in oversight of automated hiring tools. Legal commentary read it as a signal that the Department of Consumer and Worker Protection would face pressure to enforce more actively, and advised employers to expect greater scrutiny.

That is the system beginning to respond, roughly two and a half years after the law took effect.

It is worth noting what enforcement can and cannot reach. Penalties run from $500 for a first violation to $1,500 for subsequent ones, with each day of use counted separately, so sustained non-compliance can accumulate. What enforcement cannot easily do is second-guess a scope determination, because that requires knowing what a tool does inside a company that has said it does something else.

The enforcement gap and the definitional gap are the same gap.

Why this matters beyond New York

The reason to study this law closely is that it is the template.

Illinois had regulated AI video interviews since 2020 without an audit mandate. Colorado, the EU AI Act and multiple state proposals have since adopted the same basic architecture: classify a system as high-risk or in-scope, require an assessment, require disclosure.

Every one of them inherits the same structural question. Who determines scope, and what happens when the answer is "the regulated party"?

The EU AI Act defines employment as a high-risk category with obligations attached, and it too depends on classification decisions made in the first instance by providers and deployers. The mechanism differs; the dependency does not.

Null compliance is therefore not a New York problem. It is the default failure mode of any regulation that asks a party to self-identify into scope and then places the burden of disproving that determination on a regulator who cannot see inside the system.

The design that would avoid it is a registry: a requirement to declare the tools in use regardless of scope determination, so absence becomes meaningful. No jurisdiction has adopted one.

The pattern across four domains

Four articles into this series, a shape has emerged that none of them shows alone.

Medicine measures the wrong thing. 1,524 cleared devices, 1.6% citing a randomised trial. Clearance certifies resemblance to an existing product, so the regulator's question and the clinician's question are different questions.

Law measures the right thing by accident. 1,313 documented failures, not because law is worse but because an adversary reads every filing. The measurement exists as a by-product of litigation, not because anyone designed it.

Education measures the easy thing. Satisfaction at 0.93 against knowledge at 0.53. Self-report moves further than performance, and self-report is what most instruments capture.

Hiring measures nothing at all. 18 audits from 391 employers, and the ones that exist mostly pass. The regulation asks a question the regulated party gets to decide whether to answer.

The common failure is not that these systems are bad. It is that the measurement apparatus in each domain was built for something else and has been pointed at AI without being redesigned. Clearance was built for instruments. Litigation adversarialism was built for lawyers. Education instruments were built for classroom research where blinding was already impossible. Employment auditing was built on a four-fifths heuristic from the 1970s.

The domains where AI looks best are the domains measuring least, and the domain that looks worst is the only one with an adversary. That ordering should be read as a fact about instrumentation rather than about the technology, and it is the single most useful thing to carry from this series into a domain it does not cover.

What to do if you are actually evaluating a hiring tool

Six questions, and none of them is answered by an audit report.

Ask for the impact ratio on your own pipeline, not the vendor's. The vendor's audit was run on their data. Yours has a different applicant pool, a different role and a different baseline, and the ratio is a property of the combination rather than of the tool.

Ask what the tool would have to see to be biased. If it never receives a protected characteristic, ask which features correlate with one. Postcode, school, employment gaps and language patterns all do.

Ask about the four-fifths result and the underlying rates. A ratio of 0.85 between two groups selected at 3% and 3.5% is a different situation from the same ratio at 30% and 35%.

Ask what happens when it is wrong. A screening tool's false negatives are candidates who never learn they were rejected by a machine. That is the failure with no feedback loop, and it is invisible to every metric the vendor reports.

Ask whether a human can actually override it. A recruiter reviewing a ranked list is not overriding it; they are reading it in the order it was given. Meaningful override requires seeing what was filtered out.

And run the audit before deployment, not annually afterwards. The regime's cadence is a compliance artefact. Nothing stops you measuring on your own data before you rely on it.

What is unresolved

Whether the tools are actually biased. The remarkable thing about the first algorithmic accountability law in the world is that after three years it has not answered its own question. The published audits mostly pass, the unpublished ones do not exist to examine, and the population parameter is unknown.

Whether a registry would work. Requiring declaration of tools in use regardless of scope would make absence meaningful, and would also impose reporting on a very large number of ordinary software systems. Nobody has drafted a version that is both effective and proportionate.

Whether disclosure changes behaviour at all. The theory of change is that publication creates pressure. If the intended readers cannot find or parse the disclosures, the mechanism has no path to the outcome, and the research suggests they cannot.

And what happens to candidates who were screened out. They do not know it happened, cannot request the reasoning, and are not represented in any audit. The entire measurement apparatus observes the selected.

The counter-argument

Being first means being imperfect, and that is not an argument against trying. Local Law 144 is the world's first attempt at a third-party algorithmic audit regime. Discovering that scope self-determination undermines it is exactly the kind of finding that only comes from implementation, and the finding is now available to every jurisdiction drafting a successor.

The compliance figures may understate reality. The study sampled employers advertising roles, and an employer may have audited without posting where the researchers looked, or may truly not use a covered tool. Many organisations hire without any automated screening. A 4.6% posting rate does not establish a 95% non-compliance rate.

Vendors did change behaviour. An audit industry now exists, bias testing has become a procurement question, and vendors publish reports they previously would not have produced. That is a real shift, and it happened because of a law with weak enforcement, which suggests the mechanism is not purely symbolic.

And a high pass rate might just be true. Disparate impact in hiring tools has been a known liability since the well-publicised withdrawal of an early resume-screening system, vendors have had years to test for it, and it is possible that most commercial tools now clear four-fifths because their builders made sure they would. The selection-bias reading is more interesting and it is not the only reading.

The short version

New York City's Local Law 144, effective July 2023, was the world's first law requiring independent bias audits of automated hiring tools. Researchers checked 391 employers. Eighteen had posted an audit report; thirteen had posted the required transparency notice. Nearly every audit that existed reported an impact ratio above 0.8, the conventional passing threshold.

Neither figure means what it looks like. The law lets employers determine whether their own tool is in scope, and the definition turns on whether a system "substantially assists" a decision, which is a judgement the employer makes. So a missing audit cannot be distinguished from a tool declared out of scope. The researchers call this null compliance, and it makes the low number uninterpretable rather than merely damning.

The high pass rate compounds it. The published set is not a sample of tools; it is a sample of tools whose owners chose to publish. A 96% publication gap alongside a near-total pass rate among publishers is the exact pattern selection bias produces, and nothing in the data rules it out.

The law also does not require fixing anything. It mandates audit, publication and notice, not remediation. Its theory of change is that disclosure creates pressure, which requires someone to read the disclosures, and the same research found the regime of limited value to job seekers because of accessibility and usability shortcomings.

The reason this matters beyond New York is that it is the template. Colorado, the EU AI Act and multiple state proposals adopt the same architecture: classify as in-scope, assess, disclose. Every one inherits the same question of who determines scope. Null compliance is the default failure mode of any regulation that asks a party to self-identify into it, and the fix, a registry of tools in use regardless of scope, has been adopted nowhere.

Common questions

What is NYC Local Law 144? The first law anywhere to require independent bias audits of commercial algorithmic systems, effective 5 July 2023. It obliges employers using an automated employment decision tool for hiring or promotion in New York City to commission an annual independent bias audit, publish the results, and notify candidates at least ten business days before use. It is enforced by the Department of Consumer and Worker Protection, with penalties from $500 for a first violation to $1,500 for subsequent ones, each day of use counted separately.

How many employers actually comply with the AI hiring audit law? Researchers who checked 391 employers found 18 with posted audit reports and 13 with posted transparency notices, roughly 4.6% and 3.3%. Those figures cannot be read directly as non-compliance, because the law lets employers determine whether their tool falls in scope, so an absent audit is indistinguishable from a good-faith determination that the law does not apply.

What is null compliance? The condition where an absence of evidence cannot be interpreted, because the regulated party controls whether they are subject to the requirement. Under Local Law 144, an employer decides whether their tool "substantially assists" a hiring decision, so a missing audit could mean non-compliance or could mean a scope determination. It is the default failure mode of any regulation asking parties to self-identify into it.

Did the bias audits find bias? Almost none of them. Nearly every published audit reported an impact ratio above 0.8, the conventional four-fifths threshold. Two readings are available: hiring tools are mostly fine because vendors have tested for this for years, or employers who expect to fail define their tools out of scope rather than publishing a failing number. Nothing in the data distinguishes them, though the second is what the incentive structure predicts.

What is the four-fifths rule? A long-standing rule of thumb in US employment discrimination practice: if one group's selection rate falls below four-fifths of the highest group's rate, that is treated as evidence of adverse impact worth investigating. It measures outcomes rather than mechanism, it is a snapshot of the data it was run on, and it is a screening heuristic rather than a legal standard. An audit reporting 0.81 has failed to trip a heuristic, not proved fairness.

Does the law ban biased hiring tools? No. It requires an audit, publication of the result and notice to candidates. A tool with a poor impact ratio is not thereby illegal under Local Law 144, though a published audit showing adverse impact may create exposure under Title VII or the New York City Human Rights Law. The regime is a transparency mechanism whose theory of change is that disclosure creates pressure.

How does this compare with the EU AI Act? The architecture is the same: classify a system as high-risk or in-scope, require an assessment, require disclosure. The EU AI Act designates employment as high-risk with obligations attached, and it too depends on classification decisions made in the first instance by providers and deployers. The mechanism differs and the dependency on self-determined scope does not, which means it inherits the same structural question.

What should an employer ask before buying an AI hiring tool? Ask for the impact ratio on your own pipeline rather than the vendor's, since the ratio is a property of the tool and the applicant pool combined. Ask which features correlate with protected characteristics, since postcode, school, employment gaps and language patterns all do. Ask for the underlying selection rates rather than the ratio alone. Ask what happens to false negatives, who never learn a machine rejected them. Ask whether a human can see what was filtered out rather than only the ranked survivors. And run the audit before deployment rather than annually afterwards.

Learn the concepts

← All posts