Home/Blog/Do AI detectors work? Accuracy, bias, and false positives

Do AI detectors work? Accuracy, bias, and false positives

Students are accused, applicants are rejected, and articles are dismissed on the strength of a percentage produced by an AI detector. Those tools do not measure who wrote something. They measure how predictable it is, and the difference between those two things has real consequences for real people.

A student is called into a meeting because software gave her essay a score of ninety-four percent AI-generated. She wrote every word. She has no way to prove it, and the tool that accused her cannot explain itself beyond a number. This scene has played out often enough since 2023 to have generated its own genre of news story, and it rests on a widespread misunderstanding of what these tools do. AI detectors do not detect authorship. They detect statistical typicality, meaning how predictable and uniform a piece of writing is, so what they actually flag is clear, conventional, plainly written prose rather than machine-written prose, and because most text is still written by people, even a detector with impressive-sounding accuracy will produce more false accusations than true ones.

This guide explains what these tools measure, what the independent evidence says about their accuracy, why their errors fall hardest on non-native English speakers for reasons that are structural rather than accidental, the arithmetic of base rates that decides whether a detector is usable at all, why watermarking is a different proposition with different limits, and what actually works instead. The question is not whether detection is imperfect, since every test is, but whether a probabilistic signal should be used to make binary accusations that carry real consequences.

What AI detectors actually measure

A detector does not read your writing the way a teacher does. It has no knowledge of you, your history, or your process. It computes statistical properties of the text and compares them against patterns associated with machine generation. Two measures dominate.

The first is perplexity, which captures how predictable each word is given what came before. Because a language model generates text by choosing likely next words, one token at a time, its output tends to sit closer to the statistically expected choice, producing lower perplexity. Human writing is often more surprising, reaching for the less probable word.

The second is burstiness, the variation in sentence length and structure across a passage. Human prose tends to vary its rhythm, mixing short sentences with long ones and shifting construction as an argument develops. Machine text often maintains a smoother, more uniform cadence.

Put together, a detector is asking one question: does this text look statistically typical of machine generation? Notice what that question is not. It is not "did a machine write this." It is a correlation between surface statistics and a category, and the entire problem follows from the gap between the two.

Detection is not plagiarism checking

Much of the confusion comes from a false analogy with plagiarism software, which many people assume works similarly. It does not, and the difference is fundamental.

Plagiarism checking asks whether a passage matches existing sources, and it can answer that question with evidence. It compares text against a database, finds overlaps, and shows you exactly what matched and where. The claim is verifiable and the reasoning is inspectable.

AI detection asks whether text resembles machine writing statistically. There is no source to point to, no match to display, and nothing to verify. The output is a probability estimate dressed as a percentage, and no amount of examination reveals why beyond the fact that the prose was predictable. A document can be entirely original, never touching another source, and still be flagged, because originality and statistical typicality are unrelated properties. Treating a detection score like a plagiarism match, as evidence rather than as a weak signal, is the central mistake institutions have made.

What the evidence says

The independent record is not encouraging, and the most telling data point came from the industry itself. In 2023 OpenAI withdrew its own AI text classifier, reporting that it correctly identified only around a quarter of AI-written text while falsely flagging roughly one in eleven human-written passages. The company that built the generator could not build a reliable detector for it.

Independent academic evaluations have found similar difficulties. A multi-tool benchmark in 2023 found every detector tested scoring below eighty percent accuracy on diverse samples, with performance degrading sharply on short passages. A peer-reviewed 2024 study of six major commercial detectors reported baseline accuracy around forty percent, falling further when the machine text had been lightly edited by a person. Detectors have also produced conspicuous failures on historic human documents, including flagging the text of the United States Constitution as machine-written, which is a useful illustration of the underlying mechanism: formal, structured, conventional prose scores as predictable.

Institutions responded. More than a dozen universities, including several prominent research institutions, disabled their AI detection tools over accuracy and fairness concerns, while a substantial share of others continue to use them. Vendors, meanwhile, report far better figures than independent studies, and some newer detectors do claim large improvements, including near-zero false positive rates on the benchmarks where earlier tools failed badly. This is a contested area and it is worth being fair about it: detection has improved since 2023, and the strongest current tools are not the tools that were tested in the earliest critical studies. What has not changed is that vendor-reported accuracy consistently exceeds independently verified accuracy, a pattern familiar from AI benchmarks generally, and that the structural problems described below do not go away with better classifiers.

The bias problem, and why it is structural

The most serious finding concerns who gets falsely accused. A widely cited Stanford study in 2023 ran detectors over essays written by non-native English speakers and found that they were flagged as machine-written at very high rates, with an average false positive rate above sixty percent across the tools tested and one detector flagging nearly all of them, while the same detectors were close to perfect on essays by native-speaking American students. The study has real limitations, using a modest sample of relatively short essays, and at least one vendor has disputed the finding using its own larger evaluations, so the magnitude is contested.

The mechanism, however, is not in dispute, and it is what makes this more than a bug. Detectors flag low perplexity, meaning predictable word choice. Writers working in a second language tend to stay closer to standard syntax and use more common vocabulary, precisely because they are being careful. That produces exactly the statistical signature the detector was built to catch. The bias is not an accident of a particular training set that a better dataset would fix. It is a direct consequence of what the tool measures, which means it will tend to reappear in any detector built on the same principle. The same logic implicates other groups: technical and scientific writers trained to be plain and consistent, people writing in a formulaic professional register, and anyone whose style is simple and clear. This is a textbook case of the pattern described in AI bias and fairness, where a system's disparate impact follows from what it measures rather than from any intent to discriminate.

This produces an inversion worth stating plainly. Advice to write clearly, use straightforward vocabulary, and keep sentences consistent is good writing advice, and it makes you look more machine-like to a detector. The people most likely to be falsely accused are those writing carefully in a plain style, which is not a group anyone would choose to penalize.

The base rate problem

The argument that settles this is arithmetic rather than opinion, and it is the piece most discussion misses. A detector's usefulness depends not only on its accuracy but on how common machine-written text actually is in the pool being tested.

Take a hypothetical detector that is right eighty percent of the time on machine text and wrongly flags only five percent of human text, which is better than most independently measured tools. Apply it to a thousand student essays where one in ten was machine-written. It catches eighty of the hundred machine essays. It also falsely flags five percent of the nine hundred human essays, which is forty-five people. Of the hundred and twenty-five accusations produced, forty-five are innocent, more than a third.

Now suppose only one essay in twenty is machine-written. The detector catches forty of the fifty, and falsely flags about forty-eight of the nine hundred and fifty human essays. Now the majority of accusations are false. Nothing about the detector changed; only the prevalence did. This is the base rate problem, and it means a tool can be accurate in the laboratory sense and still be wrong most of the times it accuses someone, whenever the thing it looks for is uncommon. Every screening test for a rare condition faces the same arithmetic, which is why medical screening is followed by confirmatory testing rather than treated as a diagnosis. It is the same lesson that applies to misleading benchmark numbers: a headline accuracy figure tells you very little until you know what it is being applied to. AI detection is typically used with no confirmatory step at all, and the cost of the error lands entirely on the accused person, who is asked to prove a negative about their own thinking.

Why light editing defeats them

There is a further practical asymmetry that undercuts the enforcement case. Because detection depends on statistical fingerprints, changing those statistics evades it. Independent testing consistently finds accuracy falling substantially when machine output is lightly edited by a person, restructured, or run through paraphrasing tools, and a whole category of software exists specifically to adjust the perplexity and burstiness profile of text until it passes.

The consequence is uncomfortable. Someone deliberately concealing machine assistance can usually do so with modest effort, while someone writing honestly in a plain style has no similar recourse and cannot make their own writing look less predictable on demand. Detection therefore falls hardest on the people not trying to evade it, which is the opposite of what an enforcement tool should do. Hybrid work is the hardest case of all, and it is now the normal case, since drafting with assistance and then rewriting by hand produces text that no statistical method can cleanly categorize.

What detectors cannot tell you at all

Even a perfectly accurate detector would not answer the question institutions actually care about. These tools measure the statistical properties of a finished text. They have no access to process, so they cannot distinguish text a model wrote outright from text a person wrote after using a model to brainstorm, or text a person wrote and then asked a model to tidy. They cannot tell whether assistance was permitted or forbidden, disclosed or hidden, substantial or trivial. They measure correlation with a pattern, not authorship, intent, or process, and the policy questions that matter are all about the latter. A number expressing how predictable someone's prose is cannot resolve a question about how they worked.

Watermarking: a different approach with different limits

One technique deserves separating from the rest, because it is not guessing. Watermarking embeds a deliberate statistical signal into text at the moment it is generated, subtly biasing the model's word choices in a pattern that a matching detector can recognize later. Because it looks for a planted signature rather than inferring from style, it avoids the core problem: it is not penalizing you for writing plainly.

Its limits are different but real. It only works if the model that produced the text was watermarking in the first place, so it cannot cover open models, older systems, or anyone who chooses not to participate. The signal degrades under paraphrasing, translation, and heavy editing. And the absence of a watermark proves nothing about human authorship, only that no participating watermarked system left a trace. Related provenance approaches, which attach cryptographic metadata about how a file was created, have the same shape of strength and weakness, since metadata can be stripped and its absence proves nothing. These are worth pursuing and they are structurally sounder than style-based detection, but they establish origin where it was recorded rather than detecting its absence.

What actually works

If detection is unreliable, the practical question is what to do instead, and the answer is to stop looking for a technical verdict and start looking at process. Evidence of how work was produced, drafts, revision history, notes, and version records, is far more informative than a statistical guess about the finished artifact, and it is difficult to fabricate convincingly. A short conversation about a submitted piece reveals understanding in a way no classifier can, and it is the method teachers used before software existed. Where the concern is that assessment can be trivially automated, redesigning the task is more durable than policing it, since work requiring personal experience, in-class production, or engagement with specific recent material is resistant by construction rather than by surveillance.

And if a detector is used at all, its output belongs at the beginning of an inquiry rather than at the end of one. Treated as a reason to ask a question, it is a weak but non-zero signal. Treated as proof, it is a percentage being used to make an accusation it cannot support, against a person who has no way to answer it.

The short version

AI detectors do not identify authorship. They measure statistical properties of text, chiefly perplexity, which is how predictable the word choices are, and burstiness, which is how much sentence structure varies, then report how typical the text looks of machine generation. This differs fundamentally from plagiarism checking, which matches against sources and can show its evidence, whereas a detection score cannot be verified or explained. Independent evaluations have found accuracy well below vendor claims, with OpenAI withdrawing its own classifier in 2023 after it identified only about a quarter of machine text while falsely flagging human writing, and later studies finding commercial tools performing poorly on diverse and lightly edited samples. False positives fall disproportionately on non-native English speakers for a structural reason: writing carefully in standard syntax produces exactly the low-perplexity signature detectors flag. The decisive problem is base rates, since a detector that is right eighty percent of the time on machine text and falsely flags five percent of human text will accuse more innocent people than guilty ones whenever machine writing is uncommon. Light editing defeats detection, so it penalizes honest plain writers more reliably than people concealing assistance. Watermarking and provenance metadata are structurally sounder because they look for a planted signal, but only work when the generator cooperates and degrade under editing.

The idea to hold onto is that an AI detector measures how predictable your writing is, not who wrote it, so it systematically flags clear conventional prose, falls hardest on people writing carefully in a second language, and, because most writing is still human, produces mostly false accusations at realistic base rates, which makes it a weak signal that should never carry the weight of proof. If you are ever on the receiving end of one of these scores, the arithmetic above is the argument to make.

Common questions

Do AI detectors actually work? Not reliably enough to justify the way they are commonly used. Independent evaluations consistently find accuracy well below what vendors claim, with a 2023 multi-tool benchmark showing every detector scoring under eighty percent on diverse text and a 2024 peer-reviewed study of six commercial tools finding baseline accuracy around forty percent. Performance collapses on short passages and on machine text that has been lightly edited. OpenAI withdrew its own detector in 2023 after it caught only about a quarter of machine-written text while falsely flagging human writing. Detection has improved since then and newer tools claim better results, but independent verification continues to lag vendor figures.

How do AI detectors work? They analyze statistical properties of text rather than reading it for meaning. The main signals are perplexity, which measures how predictable each word is given the preceding context, and burstiness, which measures variation in sentence length and structure. Language models tend to produce text that is more predictable and more uniform than human writing, so detectors estimate how closely a passage matches that profile. The output is a probability that the text is statistically typical of machine generation, which is a different question from whether a machine actually wrote it, and that gap is the source of most detector failures.

Why do AI detectors flag human writing? Because they measure predictability rather than authorship, and plenty of human writing is highly predictable. Clear, conventional, plainly written prose, formal or technical writing, and text following a standard structure all produce the low perplexity and uniform rhythm that detectors associate with machines. This is why detectors have flagged historic documents such as the United States Constitution as machine-written. The uncomfortable implication is that ordinary good writing advice, use plain vocabulary and keep sentences consistent, makes text look more machine-like to these tools, so careful writers are among the most likely to be falsely accused.

Are AI detectors biased against non-native English speakers? The evidence points that way, though its magnitude is contested. A Stanford study in 2023 found detectors flagged essays by non-native English writers as machine-generated at very high rates while performing nearly perfectly on native-speaker essays, and at least one vendor has disputed those results using its own larger datasets, while the original study used a modest sample of short essays. The mechanism, however, is not disputed. Writers working in a second language tend to use simpler vocabulary and more standard syntax, which produces the low-perplexity signature detectors are built to flag, so the bias follows from what the tools measure rather than from a fixable dataset flaw.

What is the base rate problem in AI detection? It is the reason a detector can be accurate and still be wrong most of the times it accuses someone. Consider a detector that correctly flags eighty percent of machine text and falsely flags five percent of human text, applied to a thousand essays where one in ten is machine-written. It catches eighty real cases but also falsely flags forty-five of the nine hundred human essays, so over a third of its accusations are innocent people. If only one essay in twenty is machine-written, false accusations outnumber correct ones. Nothing about the detector changed, only how common the thing it looks for is.

Can AI detectors tell if I only used AI a little? No. Detectors analyze the statistical properties of finished text and have no access to how it was produced. They cannot distinguish text written entirely by a model from text a person wrote after brainstorming with one, or from text a person wrote and then asked a model to polish. They also cannot tell whether assistance was permitted, disclosed, or trivial. Mixed human and machine writing is the hardest case for them, and it is now the common case. Since the policy questions that matter are about process and permission, a score describing how predictable prose is cannot answer them.

What works better than AI detection? Evidence about process rather than statistical guesses about the finished text. Drafts, revision history, notes, and version records show how work developed and are difficult to fabricate convincingly. A brief conversation about a submitted piece reveals understanding more reliably than any classifier. Where the worry is that a task can be trivially automated, redesigning it, so that it requires personal experience, in-person work, or engagement with specific recent material, is more durable than policing it. Watermarking and provenance metadata are structurally sounder than style-based detection but only work when the generating system participates, and their absence proves nothing.

Sources & further reading

The primary literature behind the claims above, drawn from the concept entries this post links to, so a claim carries the same source here as it does there.

  • Jelinek et al. (1977), Perplexity — a measure of the difficulty of speech recognition tasks — where the measure comes from. Perplexity
  • Shannon (1951), Prediction and Entropy of Printed English — the floor; how predictable language actually is. Perplexity
  • Ouyang et al. (2022), Training language models to follow instructions with human feedback — the source of the term "alignment tax". Note the paper largely answers it: mixing pretraining gradients back in (PPO-ptx) removes most of the regression. :: https://arxiv.org/abs/2203.02155 Perplexity
  • Kirchenbauer et al. (2023), A Watermark for Large Language Models — the green-list method; it works. Watermarking
  • Sadasivan et al. (2023), Can AI-Generated Text be Reliably Detected? — the impossibility bound as distributions converge, and the paraphrase attack. Watermarking
  • Liang et al. (2023), GPT Detectors Are Biased Against Non-Native English Writers — the deployed harm, measured. Watermarking
  • Longpre et al. (2023), The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI — the audit; most licence tags on popular datasets are wrong or missing. Data Provenance
  • Gebru et al. (2018), Datasheets for Datasets — the template; the good sections are the ones people skip. Data Provenance

Related articles

Learn the concepts

← All posts