Home/Blog/Evaluation & evidence/AI in education: students feel twice the gain they get
AI in education: students feel twice the gain they get66 randomised trials. Certainty of every result: very low.0.93satisfaction0.53knowledge66 randomised trials. Certainty of every result: very low.
66 randomised trials. Certainty of every result: very low.

AI in education: students feel twice the gain they get

A 2026 review of 66 randomised trials found AI tutoring improved satisfaction by 0.93 and confidence by 0.91, against 0.53 for knowledge. The authors rated all three very low certainty.

TL;DR. The education evidence base is larger than most people assume and weaker than the effect sizes suggest. A 2026 systematic review of 66 randomised trials found large improvements in satisfaction (0.93) and confidence (0.91), and roughly half that for actual knowledge (0.53). The authors rated the certainty of all three as very low, mainly because of poor allocation concealment and blinding. A separate meta-analysis of 49 experiments put the pooled effect at 0.449. The finding nobody leads with: students consistently report liking it and feeling more capable at roughly twice the rate they demonstrate knowing more, and the gap between those two is the whole story.

---

A systematic review published in 2026 screened 39,783 records and included 66 randomised controlled trials covering 4,911 participants in undergraduate health professions education. It is one of the better evidence bases in applied AI.

Large language model personalised learning aids, the largest subgroup, produced:

  • Satisfaction: 0.93 standardised mean difference
  • Confidence: 0.91
  • Theoretical knowledge: 0.53

Those are substantial numbers. In education research an effect of 0.5 is respectable and 0.9 is large.

The same paper rates the certainty of every one of them as very low. Most included studies carried a high risk of bias, principally from poor allocation concealment and blinding. Heterogeneity ran to I² of 86% on the knowledge outcome, meaning the studies disagreed with each other enormously.

Both of those facts are in the same paper, and almost every summary reports the first without the second. The honest position is that AI tutoring probably helps, that the effect on how students feel is roughly twice the effect on what they know, and that the literature is not yet strong enough to say much more than that with confidence.

What the numbers say when read carefully

Three separate quantities get compressed into "AI improves learning", and separating them is most of the work.

Satisfaction measures whether students enjoyed the experience and would use it again. 0.93 is a large effect and it is the easiest thing to move, because a responsive system that never sighs at a repeated question is truly more pleasant than the alternative.

Confidence measures whether students believe they have learned. 0.91, essentially the same. It is also self-reported, and it is the outcome most vulnerable to the thing being measured.

Theoretical knowledge measures whether they can answer questions correctly afterwards. 0.53, roughly half.

The gap between the first two and the third is the finding. A system that makes students feel substantially more capable while making them somewhat more capable is producing a calibration problem, and the direction of that miscalibration matters: students overestimate what they know.

That is not a hypothetical concern. It is the same shape as automation bias in clinical settings and overconfidence in model outputs generally, and in education it has a specific consequence, because a student who believes they have understood stops studying.

The certainty rating, and why it is the most important number

The review used GRADE, the standard framework for rating how much confidence to place in a body of evidence, and returned low to very low across the board.

That rating is not a criticism of the researchers. It is a description of what the underlying studies allow, and it comes from three specific problems.

Blinding is nearly impossible. A student knows whether they are using an AI tutor. A teacher knows which class received the intervention. Neither can be blinded, which means expectancy effects contaminate every self-reported outcome, and satisfaction and confidence are entirely self-reported.

Allocation concealment is frequently poor. If the researcher assigning students to groups knows what the groups are, assignment can drift toward the result the study wants, usually without anyone intending it.

And the comparator is often weak. Many studies compare an AI-assisted condition against ordinary instruction with no additional support. That measures the value of receiving extra attention, not the value of the attention being artificial. A tutoring effect and an AI effect are different quantities, and much of this literature cannot distinguish them.

Heterogeneity, which is the quiet problem

The I² statistic estimates what fraction of the variation between studies comes from real differences rather than chance. The review reports 74% for satisfaction, 64% for confidence, and 86% for knowledge.

Above roughly 75% is conventionally described as considerable heterogeneity, and 86% means the studies are, in effect, measuring different things.

A pooled estimate across studies that disagree that much is a weak summary of a scattered picture rather than a precise estimate of a real quantity. The average of a set of numbers that range widely tells you where the middle is, not what to expect from your own deployment.

The practical consequence: an institution reading 0.53 and expecting 0.53 has misunderstood the statistic. The honest expectation is somewhere in a wide range that, on some outcomes and for some AI subtypes, includes no effect at all. The review says so directly: several subcategories showed favourable point estimates with confidence intervals that crossed zero.

What the second meta-analysis adds

A separate 2026 meta-analysis synthesised 49 controlled experiments, including both randomised trials and quasi-experimental designs, and reported a pooled effect of 0.449 with a 95% confidence interval of 0.194 to 0.704, after trim-and-fill correction for publication bias.

Three things about that number are worth extracting.

The correction was applied, which is good practice and implies the uncorrected figure was higher. Publication bias in education technology research runs in a predictable direction, because a study finding no effect is harder to publish and less likely to be funded by anyone with a product.

The interval is wide. 0.194 to 0.704 spans from a small effect to a large one. That is consistent with the heterogeneity above and it means the point estimate carries less information than its two decimal places suggest.

And it includes quasi-experimental designs, which are weaker than randomised trials. Mixing them raises the sample size and lowers the average evidentiary quality, which is a reasonable trade to make and should be stated when quoting the number.

Reading an effect size without being misled by it

Effect sizes are quoted constantly in education and understood rarely, and the confusion is exploitable. Four things worth carrying.

A standardised mean difference is measured in standard deviations, not in anything you can feel. An effect of 0.53 means the average student in the treatment group scored about half a standard deviation above the average control student. On a test where scores spread widely, that is a large real gain. On a test where everyone scores between 70 and 80, half a standard deviation is five marks.

The same effect size means different things at different baselines. Moving a struggling cohort from 40% to 55% and moving a strong cohort from 85% to 91% can produce identical standardised effects and are not the same educational event. Pooled figures across mixed populations hide this completely.

Effect sizes on short interventions are systematically larger. A six-week study measures novelty as well as learning. Effects reliably shrink as duration grows, which is why the near-total absence of long follow-up in this literature matters more than any single number in it.

And an effect size says nothing about cost. A 0.5 effect for twelve pounds per student per year and a 0.5 effect for three hundred are the same number and completely different decisions. Education research reports the first half of that ratio almost universally and the second almost never.

The practical form: an effect size is only interpretable against a baseline, a duration, a population and a cost. A number quoted without all four is a marketing figure that happens to be true.

The comparison everyone reaches for, and why it misleads

Any discussion of tutoring effects eventually invokes the finding that one-to-one tutoring produced roughly two standard deviations of improvement over conventional classroom instruction. It is among the most cited results in education research and it anchors expectations badly.

Against that anchor, 0.449 looks like failure. Against the actual literature on scalable interventions, it is a good result.

A meta-analysis of 282 randomised trials of tutoring found that high-impact tutoring remains one of the most effective academic supports available, and that maintaining its quality at scale depends heavily on the human element. Programmes that scale tend to lose effect. That is the relevant comparison class: not the ideal tutor, but the interventions a district can actually deploy to every student.

So the correct framing is that AI tutoring is competitive with other scalable interventions and not with individual expert attention, and enthusiasm calibrated to the second number will be disappointed by the first.

What is deployed versus what is studied

The gap here is wider than in medicine, and in the opposite direction.

What is studied is mostly tutoring: a system that explains, questions and adapts to a student working through material. That is where the trials are, and it is a small fraction of actual usage.

What is deployed is mostly teacher-facing. Lesson planning, generating practice questions, differentiating a worksheet for three reading levels, drafting feedback, and administrative work. Districts adopting AI are largely adopting it for the adults.

Almost none of that has been trialled. There is no meaningful randomised evidence on whether AI-assisted lesson planning improves student outcomes, because the outcome is separated from the intervention by a full term and a classroom.

This asymmetry matters when reading any claim about AI in education. The evidence base describes a use case that is not the dominant deployment, and the dominant deployment is being justified by evidence about something else.

There is a third category with a genuine evidence problem of its own: assessment. AI detection tools for student work have documented false positive rates, and those errors fall unevenly, with non-native English writers flagged disproportionately. An institution acting on a detector output is making an accusation on evidence that does not support it.

How to read a claim about AI in education

Six questions, and the first two do most of the work.

Which outcome moved? Satisfaction, confidence and knowledge are three different quantities and the first two move roughly twice as much as the third. A vendor citing engagement or satisfaction has not made a learning claim.

What was the comparator? AI-assisted instruction against no additional support measures the value of extra attention. Against equivalent human tutoring measures the value of the attention being artificial. Only the second is a claim about AI.

Was it randomised, and was allocation concealed? Quasi-experimental designs are common in this literature and weaker. Concealment is frequently poor even in the randomised ones.

What is the certainty rating? If a systematic review used GRADE, the rating is stated. Very low certainty alongside a large effect size is the normal pattern here and it should be quoted alongside the effect.

What is the heterogeneity? An I² above 75% means the studies disagree substantially and the pooled estimate is a weak guide to what you should expect.

And how long was the follow-up? Nearly all of this literature measures immediate post-test performance. Retention at three months is a different question and it is largely unstudied.

What is unresolved

Whether the confidence gap causes harm. Students feeling more capable than they are is measurable and its consequences are not. Whether inflated confidence reduces study time, whether it persists past the course, and whether it affects outcomes downstream have not been tested.

Whether effects survive at scale. The tutoring literature is clear that scaled programmes lose effect, and the AI case has not been through that yet. Most trials are small, run by motivated researchers, over short periods, which is the condition under which educational interventions look best.

What happens to skills that are not tested. Post-tests measure recall and application. The capacity to work through confusion without help, which is what a struggle with difficult material builds, is not on any of these instruments and may be precisely what a responsive tutor removes the need for.

And whether the teacher-facing deployments do anything. The dominant use has essentially no outcome evidence. It is plausible that returning hours to teachers improves instruction, and plausible that it does not, and nobody has run the study.

The counter-argument

Very low certainty is the default in education research, not a special condemnation. Blinding is impossible in almost every classroom intervention, so a GRADE rating of low or very low describes most of the field, including interventions everyone agrees work. Holding AI to a standard that reading instruction and class-size reduction also fail is not a fair test.

Satisfaction and engagement are not nothing. A student who finds a subject tolerable takes another course in it. Effects on persistence and enrolment are real educational outcomes even if they are not knowledge, and dismissing a 0.93 satisfaction effect as merely feelings undervalues something that predicts long-run attainment.

The access argument applies here as it does in medicine. For a student with no tutor, the comparator is not an expert but nothing at all. An effect of 0.5 for a student who currently receives no individual support is a substantially better deal than the same effect for a student who already has a tutor, and pooled averages hide that entirely.

And the evidence base is young. Sixty-six randomised trials in an area this new is a lot, most published since 2020, and the certainty ratings partly reflect immaturity rather than a ceiling. Judging the technology by the current literature risks judging the literature.

The short version

A 2026 systematic review of 66 randomised trials covering 4,911 participants found large-language-model learning aids improving satisfaction by 0.93, confidence by 0.91 and theoretical knowledge by 0.53. The same paper rated the certainty of all three as very low, citing poor allocation concealment and blinding, with heterogeneity reaching I² of 86% on knowledge. A separate meta-analysis of 49 experiments reported a pooled effect of 0.449, interval 0.194 to 0.704, after correcting for publication bias.

The finding nobody leads with is the gap between the first two numbers and the third. Students report liking it and feeling more capable at roughly twice the rate they demonstrate knowing more, which is a calibration problem pointing in the dangerous direction: a student who believes they have understood stops studying.

The certainty rating is the most important number in the literature and the least quoted. It is not a criticism of researchers but a description of what these studies allow. Blinding is impossible, since a student knows they are using an AI tutor. Allocation concealment is frequently poor. And the comparator is often ordinary instruction with no extra support, which measures the value of receiving attention rather than the value of that attention being artificial.

Two structural points matter more than any effect size. The heterogeneity means a pooled estimate is a weak guide to your own result, with several subcategories showing confidence intervals that include no effect at all. And what is studied is not what is deployed: the trials are about tutoring, while actual district adoption is largely teacher-facing lesson planning and administration, for which there is essentially no outcome evidence.

The honest summary: AI tutoring is competitive with other interventions that can reach every student, and it is not competitive with individual expert attention. Anyone anchored on the two-sigma tutoring result will be disappointed, and anyone comparing against what a school can actually deploy to everyone will not be.

Common questions

Does AI tutoring actually improve learning outcomes? Modestly, and the evidence is weaker than the headline numbers suggest. A 2026 review of 66 randomised trials found a standardised effect of 0.53 on theoretical knowledge, against 0.93 for satisfaction and 0.91 for confidence, with the authors rating the certainty of all three as very low. A separate meta-analysis of 49 experiments reported a pooled effect of 0.449 with an interval running from 0.194 to 0.704.

Why is the certainty of the evidence rated low? Three reasons specific to education research. Blinding is nearly impossible, since a student knows whether they are using an AI tutor and a teacher knows which class received the intervention, so expectancy effects contaminate self-reported outcomes. Allocation concealment is frequently poor. And many studies compare AI-assisted instruction against ordinary teaching with no additional support, which measures the value of extra attention rather than the value of the attention being artificial.

What is the difference between satisfaction and learning in these studies? They are separate measured outcomes and they move by different amounts. Satisfaction asks whether students enjoyed the experience, confidence asks whether they believe they learned, and knowledge tests whether they can answer correctly afterwards. The first two came in around 0.9 and the third around 0.5. Students report feeling roughly twice as improved as they demonstrate being, which is a calibration problem rather than a learning one.

How does AI tutoring compare with human tutoring? Unfavourably against individual expert attention, and competitively against interventions that can actually reach every student. The frequently cited two-standard-deviation result for one-to-one tutoring is a poor anchor, since a meta-analysis of 282 tutoring trials found that programmes lose effect as they scale and that quality at scale depends heavily on the human element. Compared with what a district can deploy to all students, an effect around 0.5 is a good result.

What does I² of 86% mean in these results? That roughly 86% of the variation between studies comes from real differences rather than chance, which is conventionally described as considerable heterogeneity. In practice it means the studies are measuring meaningfully different things, and a pooled average across them is a weak guide to what any particular institution should expect. Several subcategories in the same review had confidence intervals that included no effect at all.

Is AI actually being used for tutoring in schools? Less than the research suggests. The trials are mostly about student-facing tutoring, while district adoption is largely teacher-facing: lesson planning, generating practice questions, differentiating materials by reading level, drafting feedback and administrative work. There is essentially no randomised evidence on whether those uses improve student outcomes, because the outcome is separated from the intervention by a term and a classroom.

Are AI detection tools reliable for catching student cheating? No, and using them to make accusations is not supportable on the evidence. They have documented false positive rates and the errors fall unevenly, with non-native English writers flagged disproportionately. An institution acting on a detector output is making a serious accusation on evidence that does not carry it, and the appropriate use is as a prompt to have a conversation rather than as a finding.

What should a school ask before buying an AI education product? Which outcome moved, since satisfaction and confidence are not learning. What the comparator was, since against no extra support is a different claim from against equivalent human tutoring. Whether the study was randomised and whether allocation was concealed. What certainty rating any systematic review assigned, since very low alongside a large effect is the normal pattern here. What the heterogeneity was. And how long the follow-up ran, since nearly all of this literature measures immediate post-test performance and retention is largely unstudied.

Learn the concepts

← All posts