The flags land on lower prior attainment
Detection tools misclassify human writing in 10 to 20% of cases, and an analysis of 10,725 assessments found the flags falling disproportionately on younger students, male students and those with weaker prior results.
TL;DR. The largest independent evaluation of AI detection tools tested 14 tools across 126 documents and found human-written text flagged as machine-generated in 10 to 20% of cases. In a school of 1,000 students that is 100 to 200 false accusations a year. And an analysis of 10,725 student assessments found the flags are not evenly distributed: male students, younger students, and those with lower prior educational attainment are more likely to be flagged. The behavioural picture is not moral decline. 67% of students report using AI at least weekly and 8% believe their usage constitutes cheating, while only 18% of faculty across twelve universities felt they could clearly distinguish AI-aided learning from misconduct. One university recorded thousands of alleged cases and dismissed a substantial portion. 65% of institutions have changed assessment methods, which is the response with evidence behind it.
---
Status: one strong independent evaluation, one demographic analysis, and a survey landscape of variable quality. The Weber-Wulff evaluation and the 10,725-assessment analysis are the load-bearing sources. Survey figures on student and faculty behaviour come from mixed sources and are attributed where used. This article concerns detection and assessment policy. It is not guidance for any individual case.
---
What detection actually does
Weber-Wulff and colleagues conducted the largest independent evaluation of AI detection tools, testing 14 tools across 126 documents.
The headline finding is that false positive rates are unacceptable for the use to which the tools are put. Detection software incorrectly flagged human-written text as machine-generated in 10 to 20% of cases.
In a school of 1,000 students, that is 100 to 200 students falsely accused each year.
Independent evaluations conclude these tools are unreliable as evidence in academic misconduct cases, and the accompanying recommendation is that they should not be used as sole evidence.
Vendor-claimed accuracy across the market runs from 26% to 80%, which is a range wide enough to be uninformative, and one widely deployed tool reports its own false positive rate at around 9%.
Territory 9 established the underlying reason: perplexity-based detection measures how predictable text is, and predictable writing is produced by many people for many reasons that have nothing to do with machines.
Who the errors fall on
This is the finding that distinguishes the educational case from the general one.
An analysis of 10,725 student assessments across two cohorts, using a widely deployed detection tool, found that male students, younger students, and those with lower prior educational attainment are more likely to have their work flagged as AI-generated.
Prior attainment is the variable that matters most.
A student with weaker earlier results writes in ways a perplexity-based detector finds more predictable, for reasons including smaller vocabulary range, more conventional sentence construction, and closer adherence to taught templates. All of those are characteristics of somebody still learning to write.
Which means the tool is most likely to accuse the students least equipped to contest the accusation, and least likely to accuse the students whose writing is distinctive enough to look human.
Territory 9 found the same structure across a different axis, with non-native writers flagged at 61.3% against roughly 3% for native speakers. This is the same mechanism sorting on a different characteristic, and it is error asymmetry in a setting where the false positive carries a disciplinary process.
And the demographic breakdown exists because somebody chose to compute it. Most deployments do not, which is the unrecorded stratifier problem: an institution running detection without a demographic audit cannot know which of its students it is accusing disproportionately.
The behavioural picture is not what the framing suggests
The easiest account of this subject is moral decline, and the survey evidence does not support it.
67% of students report using AI at least weekly for assignments. 8% believe their usage constitutes cheating.
That gap is not defiance. It is an undefined rule.
Only 18% of faculty, in a survey across twelve universities in the UK, Australia and Canada, felt they could clearly distinguish AI-aided learning from outright cheating.
When the people setting the standard cannot state it, a student using a tool for brainstorming, translation, structuring or checking has no way to know which side of a line they are on, and the line differs by course, by instructor and by assignment.
One analysis puts the substantive point directly: the story is not that students discovered a new way to cheat, but that schools built assessment systems around tasks generative AI is unusually good at faking.
Which is not an excuse and is a diagnosis. A rule that two-thirds of a population breaks weekly while believing they are compliant is a rule that has not been communicated, and 62% of students in one survey said they wanted training.
What happens when detection is trusted
One documented case shows the cost of relying on the tools at scale.
An Australian university recorded thousands of alleged AI-related misconduct cases and reportedly dismissed a substantial portion of them.
Every dismissed case consumed staff time, student time, and the trust of a student who was accused and cleared.
And the aggregate consequence is worse than the individual one. An integrity process that produces a high dismissal rate teaches students that accusations are unreliable, which weakens the process for the cases that are correct.
A related dynamic is now measured. One 2026 study reports that 73% of students alter their work to evade detection, which describes students modifying legitimate writing to avoid being flagged, not only students concealing misconduct.
That is a direct cost to writing quality caused by the detection layer itself, and it is the clearest evidence that the tool has become part of the assignment.
What institutions are actually doing
The response with evidence behind it is redesign rather than detection, and 65% of institutions have changed assessment methods.
The University of Surrey has redesigned its entire curriculum and assessment policy, effective from September 2026, moving toward assessing process over outputs. Third-year civil engineering students may be asked to use AI to help design a building and then verify every output by hand calculation. English literature students continue submitting essays and may also submit drafts or revision memos.
The University of Bath is moving from a traffic-light system to a two-lane approach developed by the Association of Pacific Rim Universities, from 2026-27. Open assessments treat generative AI as optional or integral, appropriate where the tool would be expected in professional practice. Closed assessments are time-limited, invigilated and in person.
Others are piloting edit-tracking technology, requiring AI statements in submissions, or permitting use for formative work while prohibiting it in high-stakes exams.
One institution reports that faculty training reduced violations by 30% in pilots, which is a process intervention rather than a technical one.
The common feature of the measures that work is that they change what is assessed rather than trying to detect how it was produced. A viva, a portfolio, a revision memo, an invigilated exam and a hand-verified calculation are all resistant for the same reason: the artefact is not the only evidence.
Which is the constraint over classification finding this corpus reached in a different territory, arriving in education with the same shape.
The arithmetic a school should do before deploying
The false positive rate is usually quoted alone, which understates the problem, because what matters is how many of the flags are wrong rather than how many of the innocents are flagged.
Take a cohort of 1,000 submissions and assume, generously, that 15% genuinely contain undisclosed generated text.
That is 150 true cases and 850 legitimate submissions.
At a 15% false positive rate, roughly 128 legitimate submissions are flagged.
At an optimistic 70% true positive rate, roughly 105 of the real cases are flagged.
So the flags total around 233, of which 128 are wrong. More than half.
That is the number an integrity office experiences, and it explains the Australian case without any need for institutional incompetence: a process fed by a tool whose flags are majority false will dismiss a substantial portion of what it receives.
And the arithmetic gets worse as genuine misconduct falls. A school that successfully reduces misuse increases the share of its accusations that are wrong, because the true positives shrink while the false positives track the legitimate population.
Which is the perverse property worth naming. Detection performs worst in exactly the institutions where the underlying problem is least severe, and best where misuse is widespread, so a tool evaluated in a high-misuse setting will disappoint everywhere else.
None of this requires a study. It is the base rate applied to the published error figures, and it can be computed by any institution in an afternoon using its own estimated prevalence.
What the corpus keeps finding about this shape
This is the fourth subject in which classification failed and constraint held, and the four are worth setting together.
Open source contribution: detecting generated pull requests failed; requiring a reproducible test case worked, because it is free to somebody who did the work.
Prompt injection: classifying malicious input failed; bounding what an agent may do regardless of input held, because a capability the agent lacks cannot be invoked.
Content provenance: detecting synthetic media degrades as generation improves; a cryptographic chain does not.
And here: detecting generated text produces majority-false flags; assessing process, drafts, vivas and invigilated work does not need to detect anything.
The shared structure is that classification asks an unbounded question about every item forever, competing against a party who adapts, while constraint asks a bounded question once about the conditions under which work is produced.
And the shared cost is the same too. Every constraint measure is more effortful than a scan. A reproducible test case, a capability restriction, a signing chain and a viva all cost somebody real time, which is why the classification route keeps being chosen despite the record.
The corpus has now seen the pattern often enough to state it as a default rather than an observation: where a detector is proposed against an adapting counterparty, the constraint alternative is usually available, usually more expensive, and usually the one that works.
Three things this establishes
The false positive rate is incompatible with the use. 10 to 20% on the largest independent evaluation, in a process where the consequence is a misconduct allegation, and independent assessments say the tools should not be sole evidence.
The errors sort on prior attainment. Male students, younger students and those with weaker earlier results are more likely to be flagged, because a perplexity-based detector reads conventional writing as machine-like and conventional writing is what a developing writer produces.
And the rule is undefined rather than broken. 67% weekly use against 8% believing it is cheating, with 18% of faculty able to distinguish the categories, describes a communication failure rather than a discipline problem.
What it does not establish
That misconduct is not occurring. It clearly is, one figure puts undetected AI text in 12% of UK student submissions, and the redesign response exists because the problem is real.
That detection has no use. As a signal prompting a conversation, rather than as evidence, it may have value the independent evaluations do not measure.
That the demographic finding generalises. It comes from one analysis of two cohorts with one tool, and no independent replication exists.
And nothing about any individual case. Every figure here is aggregate, and an accusation is a specific claim about a specific piece of work.
What is unresolved
Whether redesigned assessment holds. The two-lane and process-based approaches begin this academic year, and no outcome data exists.
Whether detection improves. Vendor accuracy claims span 26% to 80%, no independent evaluation of current versions has been published, and the tools have changed since the largest one was conducted.
What the demographic breakdown looks like elsewhere. One analysis found the pattern; almost no institution computes it, which means most cannot know whether their own deployment sorts the same way.
And what students should be permitted to do. The 67% against 8% gap will not close until somebody states the rule, and the rule differs by discipline for defensible reasons.
What a school could do this term
Four measures, none requiring a purchase, ordered by cost.
Compute the flag precision on your own data. Take a sample of flagged submissions, adjudicate them properly, and report what share were upheld. An institution running detection without this number does not know whether its process is majority-correct, and the arithmetic above suggests many are not.
Compute the demographic breakdown. Flag rate by prior attainment, by age, by first language. The one analysis that did this found the pattern, and an institution that has not looked cannot claim its deployment is even-handed. This is a query against data already held.
State the rule per assignment rather than per institution. The 67% against 8% gap exists because policies are set at a level where they cannot be specific. A line on an assignment brief saying what use is expected, permitted and prohibited for that task costs one sentence and removes most of the ambiguity a general policy leaves.
And change one assessment. Not the curriculum, one assessment: a viva component, a required draft, a revision memo, an invigilated element. The institutions doing this at scale are redesigning everything, which is expensive and slow. A single course changing one component produces local evidence within a term.
The first two cost an afternoon and would tell an institution whether it has a problem. The second two cost more and are the ones with evidence behind them.
And the ordering matters. An institution that redesigns assessment without computing its flag precision has fixed the right problem for reasons it cannot demonstrate, which makes the change harder to defend and easier to reverse.
The counter-argument
Assessment redesign is expensive and the evidence for it is thin. Vivas, portfolios and process documentation cost far more staff time per student than marking an essay, the institutions adopting them have published no outcome data, and this article treats redesign as evidenced when it is currently a plan.
The demographic finding may reflect the underlying behaviour. If younger students and those with lower prior attainment do use AI more, then higher flag rates would be correct rather than biased, and the analysis cannot separate a detector sorting on writing style from a detector correctly identifying more frequent use. This article asserts the first reading.
A 10 to 20% false positive rate is not the operative figure if the tool is not sole evidence. Every serious guideline says detection should prompt investigation rather than determine outcome, so the relevant error rate is the one after human review, which nobody has measured and which could be much lower.
And the undefined-rule framing is generous. A student who submits generated text as their own knows what they have done regardless of institutional policy clarity, and treating a 67% usage figure as evidence of confusion rather than convenience assumes a good faith the surveys do not establish.
The short version
The largest independent evaluation of AI detection tested 14 tools across 126 documents and found human writing flagged as machine-generated in 10 to 20% of cases, which is 100 to 200 false accusations a year in a school of 1,000. Independent assessments conclude the tools should not be sole evidence in misconduct cases.
And the errors sort. An analysis of 10,725 student assessments found male students, younger students, and those with lower prior educational attainment more likely to be flagged, because a perplexity-based detector reads conventional writing as machine-like and conventional writing is what a developing writer produces. Territory 9 found the same mechanism sorting on native language at 61.3% against 3%.
The behaviour is not defiance. 67% of students use AI at least weekly and 8% think it is cheating, while 18% of faculty across twelve universities could clearly distinguish AI-aided learning from misconduct. A rule two-thirds break weekly while believing they comply is a rule nobody stated.
Trusting the tools has a documented cost. One university recorded thousands of alleged cases and dismissed a substantial portion, and 73% of students report altering their work to evade detection, which is a quality cost the detection layer created.
The response with evidence is redesign. 65% of institutions have changed assessment methods: process over outputs, two-lane open and closed assessments, drafts and revision memos, invigilated exams, hand-verified calculations. All resistant for one reason: the artefact stops being the only evidence.
Common questions
How accurate are AI detection tools? Not accurate enough for the use they are put to. The largest independent evaluation tested 14 tools across 126 documents and found human-written text incorrectly flagged as machine-generated in 10 to 20% of cases, which in a school of 1,000 students means 100 to 200 false accusations a year. Vendor-claimed accuracy across the market runs from 26% to 80%, and independent assessments conclude the tools are unreliable as sole evidence in academic misconduct cases.
Do the errors fall evenly? No. An analysis of 10,725 student assessments across two cohorts using a widely deployed detection tool found that male students, younger students, and those with lower prior educational attainment are more likely to have work flagged as AI-generated. Prior attainment is the variable that matters most: a perplexity-based detector measures how predictable text is, and a developing writer produces more conventional sentence construction, smaller vocabulary range and closer adherence to taught templates.
Why does that matter more than the overall rate? Because it means the tool is most likely to accuse the students least equipped to contest an accusation. Territory 9 found the same mechanism sorting on a different characteristic, with non-native writers flagged at 61.3% against roughly 3% for native speakers. It is also usually invisible: the demographic breakdown exists because one research team computed it, and an institution running detection without a demographic audit cannot know which of its students it accuses disproportionately.
Are students simply cheating more? The survey picture describes an undefined rule rather than defiance. 67% of students report using AI at least weekly for assignments while 8% believe their usage constitutes cheating, and only 18% of faculty across twelve universities in the UK, Australia and Canada felt they could clearly distinguish AI-aided learning from outright cheating. When the people setting the standard cannot state it, a student using a tool for brainstorming, translation or structuring has no way to know which side of a line they are on.
What happens when institutions rely on detection? One Australian university recorded thousands of alleged AI-related misconduct cases and reportedly dismissed a substantial portion. Every dismissed case consumed staff time, student time and trust, and a process with a high dismissal rate teaches students that accusations are unreliable, which weakens it for the cases that are correct. A 2026 study also reports that 73% of students alter their work to evade detection, which describes legitimate writing being modified to avoid a flag.
What are universities doing instead? 65% of institutions have changed assessment methods. The University of Surrey has redesigned its curriculum to assess process over outputs from September 2026, with engineering students verifying AI outputs by hand calculation and literature students submitting drafts or revision memos. The University of Bath is adopting a two-lane approach from 2026-27, with open assessments where AI use is optional or integral and closed assessments that are time-limited, invigilated and in person. Others are piloting edit-tracking, requiring AI statements, or permitting formative use while prohibiting it in high-stakes exams.
Why does redesign work where detection does not? Because it changes what is assessed rather than trying to determine how something was produced. A viva, a portfolio, a revision memo, an invigilated exam and a hand-verified calculation are resistant for one shared reason: the submitted artefact stops being the only evidence. That is the same finding this corpus reached in a different territory, where constraint on what a system may do held while classification of what it produced did not.
What is the strongest objection to this article? That the demographic finding may reflect underlying behaviour rather than detector bias. If younger students and those with lower prior attainment do use AI more frequently, higher flag rates would be correct, and the analysis cannot separate a detector sorting on writing style from one correctly identifying more frequent use. A second objection is that a 10 to 20% false positive rate is not the operative figure where guidelines require detection to prompt investigation rather than determine outcome, since the error rate after human review is the one that matters and nobody has measured it.
Sources
Primary documents only. Where a claim rests on a single report, the entry says so.
- AI and Academic Integrity: Why Detection Fails Structural Learning, on Weber-Wulff et al. 2023 The largest independent evaluation of AI detection tools, testing 14 tools across 126 documents and finding human-written text flagged in 10 to 20% of cases, with the conclusion that the tools should not be used as sole evidence in misconduct proceedings.
- Beyond Detection: Designing AI-Resilient Assessments with Automated Feedback arXiv:2503.23622 The analysis of 10,725 student assessments across two cohorts finding that male students, younger students and those with lower prior educational attainment are more likely to have work flagged as AI-generated.
- Are universities returning to in-person exams to combat AI cheating? Times Higher Education The institutional responses: the University of Surrey assessing process over outputs from September 2026, the University of Bath adopting a two-lane open and closed model from 2026-27, and edit-tracking pilots elsewhere.
- AI in Education by 2026: Assessment Panic, Cheating Doubts, and Better Integrity Rules Windows Forum The Australian case where thousands of alleged AI-related cases were recorded with a substantial portion dismissed, and the framing that schools built assessment around tasks generative AI is unusually good at faking.
Further reading
The primary literature behind the claims above, drawn from the concept entries this post links to, so a claim carries the same source here as it does there.
- Kleinberg et al. (2016), Inherent Trade-Offs in the Fair Determination of Risk Scores — why error rates cannot be equalised across groups while calibration holds. :: https://arxiv.org/abs/1609.05807 Error Asymmetry
- Wong et al. (2021), External Validation of a Widely Implemented Proprietary Sepsis Prediction Model — alert burden against omission, where only one side leaves a record. :: https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2781307 Error Asymmetry
- Daneshjou et al. (2022), Disparities in dermatology AI performance on a diverse curated clinical image set — a purpose-built stratified benchmark, which is the only retrospective remedy available. :: https://www.science.org/doi/10.1126/sciadv.abq6147 Unrecorded Stratifier
Related articles
- 232 studies, and 1.3% recorded skin typeA systematic review found AI detecting skin cancer at 90% accuracy across 232 studies. Almost none of those studies recorded who the patients were, and the ones that checked found performance dropping to chance.
- The sepsis model caught 7% of what clinicians missedA sepsis warning system ran at hundreds of US hospitals before anyone outside the vendor validated it. The external check found the number that matters is not the one being reported.
- Worse than chance means bias, not noisePeople identify high-quality synthetic video 24.5% of the time. A coin would do better, and the reason it beats them is the finding.
- Generation takes seconds. Debunking takes hours.Four open source projects closed their doors in one month. The cause is not bad contributions but a cost ratio that inverted, and the same maintainer who shut his bounty credits the technology with finding 100 real bugs.