Home/Blog/Evaluation & evidence/Worse than chance means bias, not noise

Worse than chance means bias, not noise

People identify high-quality synthetic video 24.5% of the time. A coin would do better, and the reason it beats them is the finding.

TL;DR. A meta-analysis of 56 studies covering 86,155 participants found people correctly identify high-quality synthetic video 24.5% of the time. A coin lands on 50%. Being reliably worse than chance is not incompetence, it is systematic bias: in a University of Florida study published February 2026, participants misclassified synthetic images as real 69% of the time, while a machine classifier reached 97% on the identical images. An iProov study of 2,000 consumers found 0.1% identified every item correctly. And the second-order effect is larger than the first. Research in the American Political Science Review in 2024 gave the first empirical confirmation of the liar's dividend: politicians who falsely label authentic evidence as fabricated successfully reduce accountability. The capability does not need to be used to do damage. Its existence is enough.

---

Status: established. Sources: a 2024 meta-analysis in Computers in Human Behavior Reports; the University of Florida study, February 2026; iProov's 2025 consumer study; the American Political Science Review 2024 liar's dividend paper; and the International AI Safety Report 2026. Several vendor-published figures appear in this literature and are labelled where used.

---

The measurement

Across 56 studies and 86,155 participants, average detection accuracy for high-quality synthetic video is 24.5%.

For images it is higher, around 62%, and across all modalities the pooled figure is 55.54%, which is barely above chance.

The video number is the one to sit with. Random guessing on a binary judgement produces 50%. A rate of 24.5% is not close to random. It is reliably, substantially worse.

And a system that is worse than chance contains information. Pure noise gives you 50%. Getting it wrong three times in four means the errors point one direction, which means something predictable is happening rather than nothing.

The direction

Participants misclassified synthetic images as real 69% of the time.

The bias is toward authenticity. When uncertain, people conclude the thing in front of them is genuine, which is the default that served for the entire period in which producing convincing fake video was expensive.

Related work finds the same shape in text. Human annotators judging scientific abstracts perform near random and tend to think all abstracts are human-written. After a five-minute conversation, participants misidentified text from a current model as human-written 77% of the time.

And exposure does not fix it. One study found no significant difference between annotators experienced with machine-generated text and those with none. Explicit warnings did not significantly improve accuracy, though they did reduce trust in the content overall, which is a different and worse outcome: less accuracy, more suspicion.

This is proxy decay operating on a human perceptual heuristic rather than a metric. Looking real was a workable proxy for being real, for as long as faking it was hard. The heuristic did not change. Its basis did.

The inversion against the previous article

The detection article established that machine detectors of machine text fail badly, and fail in a directional way that lands on non-native writers.

This literature runs the other way. In the University of Florida study, human accuracy on synthetic images was at chance, statistically indistinguishable from a coin flip, while a convolutional network on the identical images reached 97%.

Both findings are true and they are not in tension, because they concern different tasks. Detecting synthetic images is a signal-processing problem with artefacts a classifier can learn. Detecting machine-written text requires distinguishing a statistical property that human writing also exhibits.

The practical consequence is uncomfortable in both directions. For images, human judgement is unreliable and machine judgement is good, so verification should be automated. For text, machine judgement is unreliable and lands its errors on identifiable groups, so automation is the problem.

Anyone reasoning about detection in general is reasoning about two different problems, and the answer reverses between them.

The larger effect is the one nobody needs to fake

*Chesney and Citron named the liar's dividend in the California Law Review in 2019: the existence of the technology lets a wrongdoer dismiss authentic evidence as fabricated.*

It requires producing nothing. The ambient uncertainty is the mechanism. A recording that would once have settled a question now opens one.

*Research in the American Political Science Review in 2024 provided the first empirical confirmation, showing that politicians who falsely label authentic evidence as misinformation can successfully reduce accountability.*

That moves the claim from a plausible worry to a measured effect, which is the distinction this corpus exists to make, and it is why the liar's dividend rather than any individual fabrication is the substantive finding here.

The asymmetry underneath it is the same one this corpus keeps finding. Fabricating evidence takes effort, skill and risk. Denying evidence takes a sentence, and the denial gets its plausibility from work somebody else did.

Why detection tooling does not close it

Benchmark accuracy above 95% is common. Real-world performance is roughly half that, degraded by compression, adversarial pressure and generation models that move faster than the classifiers trained on them.

And the adversarial dynamic is structural. Once a generator learns which artefacts detectors flag, the next version removes them. A control that depends on detection software rests on an eroding base, which is a design property rather than a temporary state.

Which is why the serious proposals are provenance rather than detection. Signing content at capture and carrying the signature forward answers a different question, one that does not degrade as generation improves: not "does this look real" but "can this chain be verified".

Three things this establishes

Worse than chance is a finding, not a failure. A 24.5% rate means the errors are structured and point toward believing things are genuine. Training people to be more suspicious addresses the direction and not the underlying inability, and the evidence suggests it lowers trust without raising accuracy.

Detection reverses between modalities. Machines are far better than people at synthetic images and worse than useless at machine text. A policy built on either finding alone will be wrong about the other.

And the denial is cheaper than the fabrication. The liar's dividend is now empirically confirmed, costs nothing to deploy, and scales with the technology's reputation rather than its use. The damage does not require anyone to make a fake.

What it does not establish

That people are gullible. The heuristic being exploited was correct for the whole of recorded history until recently, and abandoning it has costs of its own.

That the figures are stable. Detection studies use different stimuli, different quality levels and different tasks, and pooled averages across 56 studies conceal wide variation.

That provenance solves it. Signing at capture requires adoption at the device layer, universal verification at the display layer, and does nothing about content already in circulation.

And nothing about any specific claim, recording or dispute. Every figure here is aggregate.

What is unresolved

Whether the bias can be corrected without collateral damage. Warnings reduced trust without improving accuracy, which trades one failure for another.

How the liar's dividend interacts with genuine fabrication. Both effects grow with the same underlying capability and no work separates their contributions.

Whether courts adapt. Defendants are already challenging digital evidence on these grounds, and evidentiary standards were built when recordings were hard to fake.

And whether detection stays useful at all. If real-world accuracy is half benchmark accuracy and the gap widens with each generation, the asset depreciates on a schedule nobody has estimated.

The counter-argument

Much of this literature is vendor-published. Identity verification companies, security training vendors and detection providers produce a large share of the most-quoted figures, and every one of them sells a product whose market depends on human judgement being inadequate. The peer-reviewed meta-analysis and the political science work are the load-bearing citations here for that reason, and the consumer studies are supporting rather than primary.

Laboratory detection is not field detection. Participants shown isolated clips and asked to judge authenticity are doing something nobody does in life, where content arrives with a source, a context and a plausibility prior. Real-world judgement may be considerably better than 24.5% because it is not actually a perceptual task.

The 24.5% figure may reflect stimulus selection. Studies use high-quality synthetic media by design, since testing obvious fakes would be uninformative, which means the pooled number describes the hardest cases rather than the typical ones.

And the liar's dividend has a legitimate version. Sometimes evidence genuinely is fabricated, and a public that has become harder to convince by a recording alone is not straightforwardly worse off. Appropriate scepticism and corrosive doubt are the same behaviour viewed from different sides.

The short version

Across 56 studies and 86,155 participants, people identify high-quality synthetic video 24.5% of the time. A coin manages 50%. Being reliably worse than chance means the errors are structured, and they are: participants misclassified synthetic images as real 69% of the time, and after a five-minute conversation misidentified model-generated text as human 77% of the time.

The bias runs toward believing things are genuine, which was a sound heuristic for as long as faking was expensive. The heuristic did not change; its basis did.

Exposure and warnings do not fix it. Experienced annotators performed no better than inexperienced ones, and explicit warnings failed to improve accuracy while reducing trust in the content overall.

And detection reverses between modalities. A machine classifier reached 97% on the same synthetic images where human accuracy was indistinguishable from a coin flip, while machine detectors of machine text fail badly and land their errors on non-native writers. Automate the first, do not automate the second.

The larger effect requires no fabrication at all. The liar's dividend, named in 2019 and *empirically confirmed in the American Political Science Review in 2024, shows politicians falsely labelling authentic evidence as fake can successfully reduce accountability. Making a fake takes effort and risk. Denying a real thing takes a sentence*, and it borrows its credibility from work somebody else did.

Common questions

How well do people detect synthetic media? Badly, and worse than guessing for video. A meta-analysis of 56 studies covering 86,155 participants found average detection accuracy for high-quality synthetic video at 24.5%, against 50% for a coin flip. Images are easier at around 62%, and the pooled figure across all modalities is 55.54%, barely above chance. An iProov study of 2,000 consumers found only 0.1% identified every item correctly.

How can people be worse than random? Because the errors are systematic rather than noisy. Random guessing produces 50%; being wrong three times in four means something predictable is happening. The direction is toward authenticity: participants misclassified synthetic images as real 69% of the time. When uncertain, people conclude what they are seeing is genuine, which was a sound default for the entire period in which convincing fakes were expensive to produce.

Does experience or warning help? Very little. One study found no significant difference between annotators experienced with machine-generated text and those with no prior exposure. Explicit warnings about possible machine authorship did not significantly improve detection accuracy, though they did reduce overall trust in the content, which trades a failure of accuracy for a failure of confidence rather than fixing either.

Are machines better at this? For images, dramatically. In a University of Florida study published in February 2026, human classification accuracy on synthetic images was at chance level, statistically indistinguishable from a coin flip, while a convolutional network reached 97% on the identical images. For text the position reverses: machine detectors of machine writing fail badly and concentrate their errors on non-native writers. Detection is two different problems and the answer flips between them.

What is the liar's dividend? The term legal scholars Chesney and Citron coined in 2019 for the way the mere existence of synthetic media technology lets a wrongdoer dismiss authentic evidence as fabricated. It requires producing no fake at all; the ambient uncertainty is sufficient. Research published in the American Political Science Review in 2024 gave the first empirical confirmation, showing politicians who falsely label authentic evidence as misinformation can successfully reduce accountability.

Why is that the bigger problem? Because of cost asymmetry. Fabricating convincing evidence takes effort, skill and risk. Denying real evidence takes a sentence, and the denial borrows its plausibility from work somebody else did. The effect scales with the technology's reputation rather than its actual use, so it grows even where no fabrication occurs.

Can detection software solve it? Not durably. Benchmark accuracy above 95% is common while real-world performance runs roughly half that, degraded by compression, adversarial pressure and generation models advancing faster than the classifiers trained on them. The dynamic is structural: once a generator learns which artefacts detectors flag, the next version removes them. This is why the serious proposals are provenance rather than detection, signing content at capture so the question becomes whether a chain verifies rather than whether something looks real.

How much should these specific numbers be trusted? The peer-reviewed meta-analysis and the political science work are the load-bearing citations, and the consumer studies are supporting rather than primary. A large share of the most-quoted figures in this area come from identity verification and security vendors whose market depends on human judgement being inadequate. There is also a genuine methodological objection: laboratory studies show isolated clips stripped of source and context, which is not how anyone encounters media, so field performance may be considerably better than 24.5%.

Sources

Primary documents only. Where a claim rests on a single report, the entry says so.

  1. Meta-analysis of human deepfake detection across 56 studies Computers in Human Behavior Reports, 2024 The pooled figures across 86,155 participants: 24.5% accuracy on high-quality synthetic video, around 62% on images, and 55.54% across all modalities. Named rather than linked because the journal page sits behind access controls.
  2. The Liar's Dividend: Can Politicians Claim Misinformation to Evade Accountability? American Political Science Review, 2024 The first empirical confirmation that falsely labelling authentic evidence as misinformation reduces accountability, testing a mechanism named by Chesney and Citron in the California Law Review in 2019.
  3. International AI Safety Report 2026 arXiv:2602.21012 The finding that participants misidentified model-generated text as human-written 77% of the time after a five-minute conversation, and the summary that humans detecting deepfakes often perform no better than chance.
  4. Detecting AI-Generated Text: Factors Influencing Detectability with Current Methods arXiv:2406.15583 The survey of human-as-detector studies: performance near random guessing, annotators tending to judge all text human-written, and no significant advantage for those with prior exposure.

Further reading

The primary literature behind the claims above, drawn from the concept entries this post links to, so a claim carries the same source here as it does there.

  • Raji et al. (2021), AI and the Everything in the Whole Wide World Benchmark — operationalisation standing in for the construct it was meant to represent. :: https://arxiv.org/abs/2111.15366 Proxy Decay
  • Liang et al. (2023), GPT detectors are biased against non-native English writers — a statistical signature that stopped separating the populations it was assumed to separate. :: https://arxiv.org/abs/2304.02819 Proxy Decay

Learn the concepts

← All posts