Home/Blog/Evaluation & evidence/Plus 48% with the tool, minus 17% without it

Plus 48% with the tool, minus 17% without it

Territory 12 opens on education, where the randomised evidence is unusually good and points in opposite directions depending on whether the test allows the tool.

TL;DR. A Harvard randomised crossover trial of 194 physics students found a custom GPT-4 tutor producing median learning gains more than double in-class active learning, at effect sizes of 0.73 to 1.3 standard deviations, p below 10^-8, in less time and with higher engagement. A separate randomised trial of roughly 1,000 high school students found the opposite result on a different measure: with unrestricted GPT-4 access during practice, assisted performance rose 48% and unassisted exam performance fell 17% against control. Three further experiments found assistance boosting performance while reducing persistence and performance on subsequent unassisted tasks. Both findings are real and they are not contradictory. One measures learning delivered through the tool; the other measures learning that survives its removal. Established systems measured at scale sit at 0.18 to 0.29 standard deviations.

---

Status: unusually strong randomised evidence, pointing two ways. Sources are peer-reviewed or preprint randomised trials: the Harvard crossover trial in Scientific Reports, a high school trial of roughly 1,000 students, a supervised trial from Google DeepMind across five UK schools, and large-scale trials of established tutoring systems. The disagreement is real and is the subject of this article.

---

The strong result

A randomised controlled trial assigned 194 students in an introductory physics course to learn either through a GPT-4-based AI tutor or through traditional in-class active learning, with every student experiencing both conditions in a crossover design.

Median learning gains in the AI condition were more than double those in the active learning condition. Effect sizes ran from 0.73 to 1.3 standard deviations, at p below 10^-8.

Students completed the AI lessons faster, with a median of 49 minutes, and reported higher engagement.

Two design features matter more than the headline.

The comparator was active learning, not a lecture. Active learning is the pedagogical benchmark in physics education, so the comparison is against current best practice rather than against nothing, which is the design the therapy chatbot article found missing elsewhere.

And the tutor was custom-built on the same pedagogical principles as the in-class lessons. It was not a general chatbot handed to students. The trial compares two implementations of one pedagogy, which is a much narrower and much more informative claim than the coverage suggests.

The opposite result

A separate randomised trial in a high school mathematics setting assigned roughly 1,000 students either unrestricted GPT-4 access during practice or no access.

Performance with the assistance improved by 48%.

Unassisted exam performance fell 17% relative to control.

And three further experiments found that AI assistance boosted maths and reading performance while reducing persistence and performance on subsequent unassisted tasks.

This is what the literature calls an assist-versus-test reversal, and it is consistent with a long-standing prediction from cognitive science: active cognitive engagement with material produces more durable learning than passive processing.

A student who reaches the right answer with help has not necessarily built the knowledge that produces the right answer without it, and the two are measured by different tests.

Why both are true

The trials measured different things and the difference is not subtle.

The Harvard trial measured learning delivered through the tutor, where the tutor replaces the lesson. The relevant question is whether that delivery mechanism teaches better than a classroom, and the answer was yes by a wide margin.

The high school trial measured learning that persists when the tool is removed, where the tool supplements practice rather than replacing instruction. The relevant question is whether students build durable knowledge, and the answer was no.

Those are different interventions as well as different measures. A custom tutor built on pedagogical principles, replacing a lesson, is not the same product as unrestricted model access during homework.

And the comparison a school actually faces is the second one, because unrestricted access is what students have.

Which is the measurement concentration pattern arriving in a new field within days of being named. The cheap measurement is performance during the session, which the platform already instruments. The expensive one is performance weeks later without the tool, which requires a separate assessment and a reason to run it.

The first is reported far more often. The second reverses the sign.

The scale gradient

Effect sizes fall sharply as studies get larger and systems get more established.

Established intelligent tutoring systems have demonstrated 0.18 to 0.29 standard deviations in large-scale randomised trials involving thousands of students.

Generative tutors in controlled conditions have shown 0.73 to 1.3 standard deviations.

One analysis notes that 0.18 to 0.29 is still meaningful, representing movement from roughly the 50th to the 65th percentile, and that replication at scale for the larger figures remains essential.

That gradient is the ordinary shape of educational research. Effect sizes shrink when a bespoke intervention run by its designers becomes a product run by ordinary teachers at scale, and the shrinkage is usually large.

Which means the 0.73 to 1.3 range should be read as a ceiling under favourable conditions, and the 0.18 to 0.29 range as what has survived contact with scale for an earlier generation of the technology.

The supervised case

A trial from Google DeepMind is worth separating because its design answers a different question.

165 students across five UK secondary schools used a generative model fine-tuned for pedagogy, integrated into chat-based tutoring on a mathematics platform.

Expert tutors directly supervised the model, with the remit to revise each message it drafted until they would be satisfied sending it themselves.

Tutors approved 76.4% of drafted messages with zero or minimal edits, meaning changes of one or two characters.

Students guided by the supervised model performed at least as well as students chatting with human tutors on every learning outcome measured, and were 5.5 percentage points more likely to solve the problem in question.

Three things follow.

The 76.4% figure is a measurement of the model, not the system. It says how often a trained tutor would have sent the message unchanged, which is a useful and unusual metric.

The outcome figure is a measurement of the system with a human in it. Non-inferiority to human tutors was achieved with a human tutor reviewing every message, which is a different claim from non-inferiority alone.

And the remaining 23.6% is the finding nobody quotes. Roughly one message in four needed more than trivial revision, and the trial does not establish what happens when nobody revises it.

What the evidence supports

Setting the four studies against each other produces a narrower claim than any of them makes alone.

A purpose-built tutor, designed on sound pedagogy, replacing a lesson, beats active learning in a controlled crossover. Strong evidence, small scale, favourable conditions.

A supervised model, with expert review of every message, matches human tutors. Strong evidence, small scale, and the supervision is load-bearing.

Unrestricted model access during practice improves assisted performance and damages unassisted performance. Strong evidence, larger scale, and it describes what students actually do.

Established systems at scale deliver 0.18 to 0.29 standard deviations. Strongest evidence, largest scale, and it is the only figure that has survived deployment.

The pattern is that effect and design track together. Where the intervention was designed, supervised or both, results are strong. Where the model was simply made available, the durable outcome went negative.

And the deployment most schools face is the fourth condition rather than the first, which is the gap between what the trials establish and what the coverage implies.

The four trials, sorted by design

The disagreement resolves once the studies are laid out by what was deployed and what was measured.

TrialInterventionMeasuredResult
Harvard, n=194Custom tutor replacing a lessonLearning via the tutor0.73 to 1.3 SD
DeepMind, n=165Pedagogy-tuned model, every message reviewedLearning via the systemMatches human tutors
High school, n≈1,000Unrestricted model access during practiceLearning without the toolAssisted +48%, unassisted −17%
Established systemsDeployed tutoring softwareLearning at scale0.18 to 0.29 SD

Read the second column downward and the results order themselves.

Designed and supervised at the top. Unrestricted in the third row. Deployed at scale in the fourth.

And the third column is where the sign changes. Rows one, two and four measure learning as the system delivers it. Row three is the only one that removed the tool before testing, and it is the only one that went negative.

Which produces a specific and testable claim. If the reversal is real rather than an artefact of that one design, then the top two rows would also show it if their students were tested weeks later without the tutor. Neither trial did that.

That is the single cheapest study available in this subject. Re-test the Harvard cohort unassisted, at distance. The instrument exists, the cohort exists, and the result would either dissolve the disagreement or confirm it.

Testing a prediction made two articles ago

The Territory 11 synthesis named measurement concentration and stated four forward predictions, one of which applies here directly.

The prediction was that any stage of a chain requiring new instrumentation will be under-measured relative to its importance, with the gap proportional to instrumentation cost rather than to how much the stage matters.

Education supplies an immediate test.

Performance during a session is instrumented by the platform. It costs nothing, it accrues automatically, and it is what every product dashboard reports.

Performance weeks later without the tool requires a separate assessment, a retention interval, and a reason to run it. It is expensive.

And the two point in opposite directions.

The prediction holds here, which is one instance and is worth stating because it was made before this article was written rather than after. A mechanism named retrospectively across eight clinical subjects has now been applied prospectively to a field it was not derived from.

That is not confirmation. One instance in a field selected because the disagreement was already visible is weak evidence, and the honest version is that the prediction survived its first opportunity to fail.

The stronger test is the one nobody has run. If retention testing becomes standard in AI education research and the assisted-versus-unassisted gap disappears, the mechanism was describing a temporary state of the literature rather than a structural feature of how evidence accumulates.

Three things this establishes

The sign of the effect depends on whether the test allows the tool. Assisted performance up 48% and unassisted performance down 17%, in the same trial, is not a contradiction and is not usually reported together.

Design and supervision carry the results. The strong findings come from a custom tutor built on the same pedagogy as its comparator, and from a system where expert tutors revised every message. Neither is the product a student has access to.

And effect sizes shrink with scale in the expected direction. 0.73 to 1.3 in controlled conditions against 0.18 to 0.29 for established systems across thousands of students. The larger figures have not been replicated at scale and the smaller ones have.

What it does not establish

That AI tutoring does not work. The Harvard result is strong, its comparator was best practice rather than nothing, and its design was a crossover that controls for individual differences.

That unrestricted access is always harmful. The reversal was measured on unassisted exams; whether the same students perform better on assisted tasks they will face in future is a different question with a different answer.

That 0.18 to 0.29 is the ceiling for generative systems. That range comes from an earlier generation of tutoring technology, and no large-scale generative equivalent has reported yet.

And nothing about long-term outcomes. The longest follow-up in this literature is weeks.

What is unresolved

Whether the crossover result replicates at scale. It is the single most important missing study, it is straightforward to design, and the field's own commentary says replication remains essential.

What happens without supervision. The supervised trial's 76.4% approval rate implies roughly one message in four required real revision, and no trial measures the unsupervised version.

Whether the reversal persists. Three experiments found reduced persistence on subsequent unassisted tasks, and none followed students beyond the immediate period.

And what students should be assessed on. If future work is done with assistance, an unassisted exam measures something that may no longer be the target, which is a curriculum question rather than a research one and nobody has settled it.

The counter-argument

Comparing the two headline results is comparing different interventions and calling it a disagreement. A custom tutor replacing a lesson and unrestricted model access during homework are not the same treatment, so finding different effects is expected rather than informative, and this article builds a framing on a contrast that dissolves on inspection.

The unassisted exam may be the wrong test. If professional and academic work will be done with AI assistance available, then measuring performance without it assesses a skill the curriculum may be about to stop valuing, and treating the 17% drop as unambiguous harm assumes the assessment stays fixed.

The scale gradient argument proves too much. Every educational intervention shows shrinking effect sizes at scale, so citing it against generative tutoring applies a discount that would equally discount the comparator, and the 0.18 to 0.29 figure comes from a technology generation with different capabilities.

And the supervision point may be overstated. A 76.4% approval rate with zero or minimal edits is high, the remaining messages were revised rather than rejected, and inferring that the unsupervised system would fail is an inference the trial explicitly does not support.

The short version

A Harvard randomised crossover trial of 194 physics students found a custom GPT-4 tutor producing median learning gains more than double in-class active learning, at 0.73 to 1.3 standard deviations, p below 10^-8, in less time and with higher engagement. The comparator was best practice, and the tutor was built on the same pedagogy as the lessons it replaced.

A randomised trial of roughly 1,000 high school students found the opposite on a different measure. With unrestricted GPT-4 access during practice, assisted performance rose 48% and unassisted exam performance fell 17% against control, with three further experiments finding reduced persistence on subsequent unassisted tasks.

Both are real. One measures learning delivered through the tool; the other measures learning that survives its removal.

A supervised trial across five UK schools found a pedagogy-tuned model matching human tutors, with expert tutors approving 76.4% of drafted messages with zero or minimal edits and students 5.5 percentage points more likely to solve the problem. The supervision was the design, and the remaining 23.6% is unexamined.

And established tutoring systems at scale deliver 0.18 to 0.29 standard deviations, which is the only figure in this subject that has survived deployment to thousands of students.

The pattern is that design and supervision carry the results. Where the intervention was built or reviewed, outcomes are strong. Where the model was simply made available, the durable outcome went negative, and that is the condition most schools are actually in.

Common questions

Does AI tutoring improve learning? On the strongest single trial, substantially. A randomised crossover trial of 194 introductory physics students found a custom GPT-4 tutor producing median learning gains more than double those of in-class active learning, with effect sizes from 0.73 to 1.3 standard deviations at p below 10^-8, in less time and with higher reported engagement. Two design features matter: the comparator was active learning rather than a lecture, so the comparison was against best practice, and the tutor was purpose-built on the same pedagogical principles as the lessons it replaced.

Why do other trials find harm? Because they measure a different thing. A randomised trial of roughly 1,000 high school students found that unrestricted GPT-4 access during practice improved assisted performance by 48% while unassisted exam performance fell 17% relative to control. Three further experiments found assistance boosting maths and reading performance while reducing persistence and performance on subsequent unassisted tasks. This assist-versus-test reversal is consistent with the cognitive science prediction that active engagement with material produces more durable learning than passive processing.

Are the two findings contradictory? No. One measures learning delivered through the tutor, where the tutor replaces the lesson and the question is whether that delivery teaches better than a classroom. The other measures learning that persists when the tool is removed, where the model supplements practice and the question is whether durable knowledge was built. They are also different interventions: a custom tutor built on pedagogical principles is not the same product as unrestricted model access during homework.

Which condition describes real schools? The second. Unrestricted access is what students have, and the custom-built supervised tutors that produced the strong results are research instruments rather than deployed products. That gap between the studied intervention and the available one is the main reason coverage of this literature overstates what it establishes.

What did the supervised trial show? A trial across five UK secondary schools with 165 students used a pedagogy-tuned generative model in chat-based mathematics tutoring, with expert tutors revising each drafted message until they would be satisfied sending it themselves. Tutors approved 76.4% of messages with zero or minimal edits, and students performed at least as well as those chatting with human tutors on every outcome measured, being 5.5 percentage points more likely to solve the problem. The supervision is load-bearing: this establishes non-inferiority for a system with a human reviewing every message, and the roughly one message in four requiring real revision is unexamined.

How big are the effects at scale? Smaller. Established intelligent tutoring systems have demonstrated 0.18 to 0.29 standard deviations in large-scale randomised trials involving thousands of students, which one analysis notes is still meaningful and represents movement from roughly the 50th to the 65th percentile. The 0.73 to 1.3 figures come from controlled conditions with bespoke systems, and replication at scale remains outstanding. Effect sizes shrinking as an intervention moves from its designers to ordinary deployment is the ordinary shape of educational research.

What should students be assessed on? Unresolved, and it is a curriculum question rather than a research one. If future academic and professional work is done with AI assistance available, an unassisted exam measures a skill whose value may be changing, so reading a 17% unassisted drop as unambiguous harm assumes the assessment stays fixed. That assumption is exactly what is being contested and no trial can settle it.

What is the strongest objection to this article? That comparing the two headline results compares different interventions and calls it a disagreement. A custom tutor replacing a lesson and unrestricted model access during homework are not the same treatment, so different effects are expected rather than illuminating. The defence is that the coverage of both results treats them as evidence about the same thing, which is the error the comparison exists to expose.

Sources

Primary documents only. Where a claim rests on a single report, the entry says so.

  1. AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting Scientific Reports, 2025 The Harvard crossover trial: 194 introductory physics students, a custom GPT-4 tutor against in-class active learning, with students learning significantly more in less time and reporting higher engagement and motivation.
  2. Faster Completion, Less Learning: Generative AI Reduced Study Time on Math Problems and the Knowledge They Build arXiv:2605.21629 The assist-versus-test reversal: a randomised trial of roughly 1,000 high school students where assisted performance improved 48% while unassisted exam performance fell 17% against control, and three further experiments finding reduced persistence on subsequent unassisted tasks.
  3. AI tutoring can safely and effectively support learning Google DeepMind, LearnLM, November 2025 The supervised trial: 165 students across five UK secondary schools, expert tutors revising every drafted message, 76.4% approved with zero or minimal edits, and students performing at least as well as with human tutors while being 5.5 percentage points more likely to solve the problem.
  4. The Algorithmic Turn: The Emerging Evidence On AI Tutoring Carl Hendrick The scale gradient: established systems at 0.18 to 0.29 standard deviations in large-scale trials involving thousands of students against 0.73 to 1.3 in controlled generative conditions, with the note that replication at scale remains essential.

Further reading

The primary literature behind the claims above, drawn from the concept entries this post links to, so a claim carries the same source here as it does there.

  • Wong et al. (2021), External Validation of a Widely Implemented Proprietary Sepsis Prediction Model — discrimination measured widely, and the population that mattered evaluated only when somebody chose to. :: https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2781307 Measurement Concentration
  • Raji et al. (2021), AI and the Everything in the Whole Wide World Benchmark — how the availability of a benchmark shapes what a field concludes it has measured. :: https://arxiv.org/abs/2111.15366 Measurement Concentration

Learn the concepts

← All posts