Home/Blog/Evaluation & evidence/The grader rewards what the detector flags

The grader rewards what the detector flags

AI graders assign higher scores to text with lower perplexity. AI detectors flag text with lower perplexity. The same institution runs both, on the same signal, with opposite consequences for the student.

TL;DR. Published research finds that GPT-4 assigns significantly higher scores to texts with lower perplexity, and that LLM-based essay scoring rates model-generated text more favourably than human-written text. The previous article established that perplexity-based detectors flag exactly the same property. The same institution can therefore run a grader that rewards a signal and a detector that punishes it. On accuracy, the picture depends entirely on the task: 95 to 99% on short answer, 85 to 92% on rubric-based essays against human inter-rater reliability of 80 to 90%, 75 to 85% on holistic open writing, and 65 to 78% on non-standard English. The most rigorous single study, on 463 Master's responses, found exact agreement of 30% and adjacent agreement of 45%, with a 191-student classroom pilot finding no significant correlation at all.

---

Status: a large and uneven literature, with the strongest studies giving the weakest results. Sources are peer-reviewed papers and preprints. The accuracy bands come from a vendor-compiled review and are labelled as such, while the individual studies with stated samples are the load-bearing ones.

---

The accuracy figures, sorted by task

A vendor-compiled review of 2024 and 2025 research reports agreement of 85 to 92% on rubric-based essay grading, against human inter-rater reliability typically running 80 to 90%.

By task type it gives 95 to 99% for multiple choice and short answer, 85 to 92% for rubric-based essays, 75 to 85% for open-ended holistic writing, and 65 to 78% for English language learners and non-standard English.

That source sells grading software, and the numbers are consistent with the peer-reviewed range while sitting at its optimistic end.

The bottom row is the one to hold. Accuracy on non-standard English is roughly thirty points below accuracy on short answer, and it is the population where a marking error costs the most.

Peer-reviewed work broadly agrees on the shape. State-of-the-art models achieve quadratic weighted kappa of 0.75 to 0.86 on K-12 essays, with transformer models above 0.80 in cross-prompt evaluation. Short-answer grading falls to 0.44 to 0.72 depending on complexity.

Where the rigorous studies land

The pattern this corpus keeps finding holds here: the more carefully a study is designed, the smaller the result.

*Flodén's 2025 study in the British Educational Research Journal evaluated grading of 463 Master's-level exam responses. 70% of AI grades fell within 10% of teacher scores. Exact grade agreement was 30% and adjacent agreement 45%.*

Seventy percent within ten percent sounds strong. Thirty percent exact agreement does not, and both describe the same data.

A 191-student classroom pilot found no significant correlation between AI and human scores at all, with the model grading more conservatively.

And configuration dominates. Fine-tuned models achieved kappa of 0.613 to 0.859 where zero-shot approaches on the same task achieved 0.023 to 0.327. Zero-shot short-answer grading produced negative predictive values, meaning performance worse than simply assigning every student the mean human score.

A grader that performs worse than a constant is not a weak grader. It is a system that has not been configured, and the gap between the configured and unconfigured versions is larger than the gap between the configured version and a human.

The bias that connects to the previous article

This is the finding that makes the subject more than an accuracy question.

Published work demonstrates quantitatively that GPT-4 assigns significantly higher scores to texts with lower perplexity, which is described as a self-preference bias.

And separate work reports that LLM-based automated essay scoring rates model-generated text more favourably than human-written text.

The previous article established that perplexity-based detectors flag text for being predictable, and that the flags fall disproportionately on younger students and those with lower prior educational attainment, because a developing writer produces more conventional prose.

So the two systems read the same signal and act on it in opposite directions.

A predictable essay scores higher from an AI grader and is more likely to be flagged by an AI detector.

An institution running both is telling a student that the same property of their writing is evidence of quality and evidence of misconduct, and neither tool is malfunctioning. Each is doing exactly what it was built to do with the statistic it was given.

This is not a hypothetical configuration. Detection and automated marking are frequently procured by the same institution, sometimes from the same vendor, and no published work examines what happens to a submission that passes through both.

The other biases, and their directions

Worth listing because they do not point the same way and cannot all be corrected by a single adjustment.

Central tendency. One study found the model producing more medium scores and fewer extreme scores than human raters, a pattern reported across multiple studies. The effect is to compress the distribution, which harms the strongest and weakest work simultaneously.

Length effects run opposite to human raters. Models assign higher scores to short or underdeveloped essays and lower scores to longer essays containing minor grammatical or spelling errors. Human raters award higher scores to longer essays.

Systematic offset. One field experiment measured a grader running 4.9 points above human raters, with intraclass correlation of 0.56.

And self-preference on generated text, as above.

The offsets are correctable and the reordering is not. A systematic 4.9-point bias can be subtracted. A grader that ranks a short weak essay above a long good one has changed the order, and no calibration fixes that.

The comparator nobody states

The same field experiment reported human-human agreement at an intraclass correlation of 0.47, against the model's 0.56.

Which means the model agreed with humans more than humans agreed with each other, on that task.

The authors' reading is that this indicates rubric ambiguity rather than an LLM-specific failure, and it is the correct reading.

And it reframes the whole accuracy discussion. Comparator choice determines what an agreement figure means. A grader compared to a single human marker is being compared to something with its own substantial unreliability, and most published agreement figures do not report the human-human baseline alongside.

That omission runs in the field's favour and against it in different studies, which is why it matters: an 85% agreement figure could describe a good grader or a bad rubric, and without the human-human number a reader cannot tell which.

The national-scale case

One study is worth separating because it is the only one operating at the scale where this is actually being deployed.

Two full national cohorts of trial school-leaving essay exams from Estonia were scored using an operationalised curriculum rubric, comparing model and statistical NLP assessments against human panel scores.

Automated scoring achieved performance comparable to human raters and tended to fall within the human scoring range.

Three features distinguish it. The rubric was the official curriculum rubric, operationalised rather than invented. The design was explicitly human-in-the-loop. And the study evaluated bias, prompt injection risk, and language models as essay writers, in the same paper.

That last item is unusual and worth crediting. A study of automated grading that also examines what happens when the essays are themselves generated is asking the question the deployment will actually face, which almost nothing else in this literature does.

Its conclusion is correspondingly narrow: a principled, rubric-driven, human-in-the-loop pipeline is viable for high-stakes writing assessment. Every clause in that sentence is load-bearing.

The two tools, on one submission

Tracing a single essay through both systems makes the interaction concrete.

Property of the writingThe grader doesThe detector does
Low perplexity, conventional proseScores it higherFlags it
Short and underdevelopedScores it higherNeutral
Long with minor errorsScores it lowerNeutral
Distinctive and idiosyncraticScores it lowerPasses it
Generated by a modelScores it higherFlags it

Rows one and five are the ones that matter.

A student who writes conventionally, because they are still learning, gets a better mark and an accusation.

A student who used a model gets a better mark and an accusation.

Those two students are indistinguishable to both systems, which is the precise reason detection fails and the precise reason the grader's bias is not correctable by calibration: the tools cannot separate the cases because the signal does not separate them.

And row four is the quiet one. A distinctive writer is scored down and passed through, so the system that flags the developing writer also under-rewards the strong one, compressing the distribution from both ends.

None of this requires either tool to be badly built. A perplexity-based grader and a perplexity-based detector are each doing something defensible with the statistic available. What is indefensible is running both without anyone having drawn this table.

What an institution could do before term

Three checks, none requiring a purchase, ordered by cost.

Score a sample with both tools and cross-tabulate. Take fifty submissions already marked by a human. Run the grader and the detector and plot grade against flag. If the flagged submissions cluster at the top of the AI-assigned grade distribution, the interaction described here is present in your deployment. This is an afternoon.

Report the human-human baseline alongside any agreement figure. Double-mark a sample and compute inter-rater agreement before comparing anything to a model. An institution quoting 85% agreement without knowing its own markers agree at 70% has not measured what it thinks.

And check the non-standard English band locally. The published figure of 65 to 78% is the lowest in the literature and the highest-stakes. Whether it holds in a specific cohort with a specific rubric is answerable from work already marked.

The ordering matters again. An institution that adopts automated marking without knowing its own inter-rater agreement has replaced an unmeasured process with a differently unmeasured one, and gained a vendor.

Three things this establishes

A grader and a detector can read one signal in opposite directions. Lower perplexity earns a higher grade and attracts a flag. Both tools are working correctly, and no published study examines a submission passing through both.

Configuration matters more than model choice. Kappa of 0.613 to 0.859 fine-tuned against 0.023 to 0.327 zero-shot, with zero-shot short-answer grading performing worse than assigning every student the mean. The distance between configured and unconfigured exceeds the distance between configured and human.

And the human baseline is missing from most reports. One study measured human-human agreement at 0.47 against a model's 0.56. An agreement figure without the human-human comparator cannot distinguish a good grader from an ambiguous rubric.

What it does not establish

That AI grading does not work. The national-cohort study found performance within the human range using an official rubric and a human-in-the-loop design, and short-answer accuracy is high.

That the perplexity finding is settled. It comes from specific studies of specific models, and whether it holds across current systems and configurations is unexamined.

That the accuracy bands are reliable. The task-by-task figures come from a vendor-compiled review, and the peer-reviewed spread is wider.

And nothing about any individual grade. Every figure here is aggregate, and a mark is a specific judgement about specific work.

What is unresolved

What happens to a submission that passes through both systems. The most consequential question in this article and nobody has studied it.

Whether the perplexity preference persists. It was demonstrated on particular models, and no series tracks it across versions.

What the human baseline is, generally. Reported once at 0.47 and omitted from most agreement studies, which makes the field's headline figures hard to interpret.

And whether feedback improves writing. One paper names this as an important direction and notes it is largely unexamined: almost all research measures whether the score matches a human's, not whether the student's next essay is better.

The counter-argument

The grader-detector contradiction is constructed rather than observed. The two findings come from different studies of different systems, no institution has been shown running both on the same submission, and this article builds its central claim on a mechanism nobody has documented in practice. That is stated in the article and it remains the weakest link.

Human marking is worse than this framing allows. An intraclass correlation of 0.47 between human raters is poor, human markers exhibit fatigue, ordering and halo effects, and a consistent model with a correctable offset may be preferable to an inconsistent human, which the accuracy comparison obscures by treating human scores as ground truth.

Zero-shot performance is a strawman. Nobody deploys an unconfigured grader in a high-stakes setting, so citing negative predictive values from zero-shot short-answer grading describes a configuration no serious deployment uses, and the fine-tuned figures are the relevant ones.

And the non-standard English figure may reflect the rubric. If a rubric rewards conventional academic register, then lower scores for non-standard English are the rubric operating as designed, which is an argument about curriculum rather than about the grader, and this article treats it as a technical failure.

The short version

Published work finds GPT-4 assigning significantly higher scores to texts with lower perplexity, and LLM-based essay scoring rating model-generated text more favourably than human-written text. Perplexity-based detectors flag exactly that property, and the flags fall on younger students and those with lower prior attainment.

So one institution can run a grader that rewards a signal and a detector that punishes it, with neither tool malfunctioning, and no published study follows a submission through both.

Accuracy depends entirely on the task: 95 to 99% short answer, 85 to 92% rubric-based essays against human inter-rater reliability of 80 to 90%, 75 to 85% holistic writing, and 65 to 78% for non-standard English, which is the lowest and the highest-stakes.

The rigorous studies are the least flattering. 463 Master's responses gave 70% within 10% of teacher scores, 30% exact agreement and 45% adjacent. A 191-student pilot found no significant correlation.

Configuration dominates. Fine-tuned kappa of 0.613 to 0.859 against zero-shot 0.023 to 0.327, with zero-shot short-answer grading performing worse than assigning every student the mean human score.

And the human baseline is usually absent. One study measured human-human agreement at 0.47 against a model's 0.56, with the authors reading it as rubric ambiguity rather than model failure. Without that number an agreement figure cannot distinguish a good grader from a bad rubric.

Common questions

How accurate is AI grading? Entirely dependent on task. A vendor-compiled review of 2024 to 2025 research reports 95 to 99% on multiple choice and short answer, 85 to 92% on rubric-based essay grading against human inter-rater reliability typically of 80 to 90%, 75 to 85% on open-ended holistic writing, and 65 to 78% for English language learners and non-standard English. Peer-reviewed work broadly agrees on the shape, with quadratic weighted kappa of 0.75 to 0.86 on K-12 essays and 0.44 to 0.72 on short-answer grading depending on complexity.

Why do the rigorous studies give lower numbers? Because the headline figures compress a distribution that individual studies report in detail. A 2025 study in the British Educational Research Journal evaluated grading of 463 Master's-level responses and found 70% of AI grades within 10% of teacher scores, with exact grade agreement at 30% and adjacent agreement at 45%. A 191-student classroom pilot found no significant correlation between AI and human scores. Seventy percent within ten percent and thirty percent exact agreement describe the same data.

What is the connection to AI detection? Both act on the same statistical property in opposite directions. Published work demonstrates that GPT-4 assigns significantly higher scores to texts with lower perplexity, described as a self-preference bias, and that LLM-based essay scoring rates model-generated text more favourably than human-written text. Perplexity-based detectors flag text for being predictable, with the flags falling disproportionately on younger students and those with lower prior attainment. A predictable essay therefore scores higher from an AI grader and is more likely to be flagged by an AI detector, with neither tool malfunctioning.

Has anyone studied a submission passing through both? No, and it is the most consequential gap in this subject. Detection and automated marking are frequently procured by the same institution, sometimes from the same vendor, and no published work examines the interaction. This article's central claim rests on combining findings from separate studies, which is stated as its weakest link.

What other biases are documented? Several, pointing in different directions. Central tendency, where models produce more medium and fewer extreme scores, compressing the distribution and harming the strongest and weakest work simultaneously. Length effects that run opposite to human raters, with models scoring short or underdeveloped essays higher and longer essays with minor errors lower, while humans reward length. And a systematic offset, measured in one field experiment at 4.9 points above human raters. Offsets are correctable by subtraction; reordering is not.

How does configuration affect it? More than model choice does. Fine-tuned models achieved kappa of 0.613 to 0.859 where zero-shot approaches on the same task achieved 0.023 to 0.327, and zero-shot short-answer grading produced negative predictive values, meaning performance worse than simply assigning every student the mean human score. The distance between a configured and an unconfigured grader exceeds the distance between a configured grader and a human.

Are human markers reliable? Less than the comparison usually assumes. One field experiment measured human-human agreement at an intraclass correlation of 0.47 against the model's 0.56, meaning the model agreed with humans more than humans agreed with each other on that task, which the authors read as rubric ambiguity rather than model failure. Most published agreement studies omit the human-human baseline, which makes their headline figures hard to interpret: an 85% agreement figure could describe a good grader or a bad rubric.

What is the strongest evidence for deployment? A study of two full national cohorts of trial school-leaving essay exams from Estonia, scoring against an operationalised official curriculum rubric and comparing model and statistical NLP assessments to human panel scores. Automated scoring achieved performance comparable to human raters and fell within the human scoring range. The study also evaluated bias, prompt injection risk and language models as essay writers, which almost nothing else in this literature does. Its conclusion is deliberately narrow: a principled, rubric-driven, human-in-the-loop pipeline is viable for high-stakes writing assessment.

Sources

Primary documents only. Where a claim rests on a single report, the entry says so.

  1. Creating and Evaluating K-12 GenAI Assessment Graders Through Context Engineering arXiv:2606.12422 The literature synthesis: quadratic weighted kappa of 0.75 to 0.86 on K-12 essays, the Flodén 2025 study of 463 Master's responses with 30% exact and 45% adjacent agreement, fine-tuned kappa of 0.613 to 0.859 against zero-shot 0.023 to 0.327, and zero-shot short-answer grading producing negative predictive values.
  2. Does AI Homogenize Student Thinking? A Multi-Dimensional Analysis of Structural Convergence in AI-Augmented Essays arXiv:2603.21228 The bias literature: the demonstration that GPT-4 assigns significantly higher scores to lower-perplexity texts as a self-preference bias, and the finding that automated essay scoring rates model-generated text more favourably than human-written text.
  3. LLMs Do Not Grade Essays Like Humans arXiv:2603.23714 The out-of-the-box grading behaviour: weak agreement varying with essay characteristics, higher scores for short or underdeveloped essays, lower scores for longer essays with minor errors, and the observation that whether feedback improves student writing is largely unexamined.
  4. Machine-Assisted Grading of Nationwide School-Leaving Essay Exams with LLMs and Statistical NLP arXiv:2601.16314 The national-scale case: two full Estonian cohorts scored against an operationalised official curriculum rubric, performance comparable to human raters and within the human range, with bias, prompt injection risk and models as essay writers evaluated in the same study.

Further reading

The primary literature behind the claims above, drawn from the concept entries this post links to, so a claim carries the same source here as it does there.

  • Heinz et al. (2025), Randomized Trial of a Generative AI Chatbot for Mental Health Treatment, and the published response identifying the waitlist control among three limitations. :: https://ai.nejm.org/doi/full/10.1056/AIoa2400802 Comparator Choice
  • Wong et al. (2021), External Validation of a Widely Implemented Proprietary Sepsis Prediction Model — performance against the population that mattered rather than against overall discrimination. :: https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2781307 Comparator Choice

Learn the concepts

← All posts