Lesson planning is 60 to 99% of the usage
Teachers report saving between 2.9 and 14 hours a week with AI, every figure is self-reported, and the platform data shows they use it for planning rather than for the admin the case rests on.
TL;DR. A RAND survey of 4,200 K-12 teachers in January 2026 found 68% using AI at least weekly, up from 29% a year earlier. Reported time savings range from 2.9 hours a week for monthly users to 5.9 for weekly users in one survey and 10 to 14 hours in another. Every one of those figures is self-reported, and no education equivalent of repository telemetry exists to check them. RAND's own caveat is the important line: 62% of teachers said the time saved was partially offset by time spent reviewing outputs. And platform data undercuts the framing. Across 12 districts and 4,644 teachers exchanging 412,304 messages, lesson planning accounted for 60 to 99% of usage in every district. Teachers are using it for the professional core of the job, not the paperwork the workload case is built on.
---
Status: high-quality adoption data, self-reported outcomes, and one useful telemetry source. The RAND and Gallup surveys state their samples. The platform report covers 12 districts with logged usage rather than recall. The distinction between logged and reported figures is the article's subject and is marked throughout.
---
The problem the tools are aimed at
Teachers in England work an average of 50.3 hours a week against a 37.5-hour standard, per the Department for Education's 2024 workload survey.
The OECD's Teaching and Learning International Survey, covering teachers in 48 countries, found an average of 40% of working hours spent on tasks other than direct instruction.
And the burnout literature ranks the predictors. Workload and administrative burden is the strongest. Role ambiguity and conflicting demands is second. Lack of autonomy is third. Student behaviour, which dominates popular accounts, ranks fourth.
That ordering is the most useful fact in this subject and it cuts two ways.
It supports the intervention: if administrative burden is the strongest predictor of attrition, reducing it is the highest-leverage available change.
And it bounds the intervention: the second and third predictors are things a tool cannot address. AI cannot clarify contradictory expectations from administrators, standardise inconsistent evaluation criteria, or return decision-making authority to a teacher who has none.
One analysis states the limit precisely: making unfulfilling work more efficient is not the same as making it meaningful.
The adoption numbers, which are solid
RAND surveyed 4,200 K-12 teachers across the United States in January 2026 and found 68% using AI tools at least weekly in their professional practice, up from 29% in January 2025.
A Gallup and Walton Family Foundation survey found 60% of teachers using AI over a school year, with 30% weekly.
And logged platform data shows the trajectory directly. Across 12 participating districts, 1,438 teachers used one assistant in September, 2,891 by the end of December, and 4,644 by 31 March 2026, a 61% increase, exchanging over 412,304 messages.
Active users averaged 12 sessions and 3.7 hours on the platform, at about 18 minutes per session.
Adoption is not in dispute. More than doubling weekly use in twelve months is a genuine finding across independent instruments, and the logged figures corroborate the survey ones, which is unusual in this corpus.
The time savings, which are not
Reported savings vary by a factor of five across sources.
Gallup and Walton report weekly users saving 5.9 hours a week and monthly users 2.9 hours, framed as roughly six weeks a year.
RAND reports 10 to 14 hours a week for teachers using AI across multiple categories.
Case study material reports up to 5 hours.
Every one of these is a teacher's estimate, and one methodology note is explicit that savings were estimated by the half-hour on a task-by-task basis by teachers who said AI affected their time on that task, then summed across non-overlapping tasks.
That is a reasonable survey design and it is recall arithmetic, which is the weakest form of the measure this corpus has documented producing systematically favourable results.
The self-report gap has three instances in this corpus, all pointing one way: developers estimating 20% faster and measuring 19% slower, executives at 97% benefiting against 29% organisational return, and developers reporting improved code quality against telemetry showing refactoring collapsing.
In each of those cases a second instrument existed. Here there is none. No education equivalent of repository telemetry measures what teachers actually did with the hours, so the corpus can flag the pattern and cannot resolve it.
The caveat that should be the headline
RAND's researchers noted that 62% of teachers reported the time saved was partially offset by time spent reviewing AI outputs.
Nearly two thirds, in the largest survey in the subject, reporting that the saving is partly illusory.
And this is the third domain in which the same offset appears. Ambient clinical scribes require mandatory physician review, which is stated to partially offset the capture savings. Senior engineers report 20 to 35% more code review time where colleagues lean on assistants.
Three fields, one mechanism: generation is fast, and the checking that follows lands somewhere and is not counted in the headline.
What distinguishes education is who absorbs it. In clinical documentation the reviewer is the same person who saved the time. In software the reviewer is frequently a more senior colleague. In teaching, the reviewer is the same teacher, which means the 5.9-hour figure and the 62% offset describe the same person's week, and the honest net figure is somewhere below the reported one by an amount nobody has measured.
What the platform data actually shows
This is the finding that complicates the workload case, and it comes from logs rather than recall.
Lesson planning features made up 60 to 99% of usage in every participating district. Conversations with the assistant accounted for the large majority of platform activity, with frequent users concentrating even more heavily on it and lighter users exploring a broader mix.
Survey data agrees on the ranking. Gallup's most common applications are research and content gathering at 44%, creating lesson plans at 38%, summarising at 38%, and generating classroom materials at 37%.
Grading, one-on-one instruction and student data analysis are consistently the least common uses.
Which means teachers are not primarily automating paperwork. They are using it for lesson design, which is the creative and professional core of the job and the thing the workload argument promises to free them up for.
Two readings are available and the evidence does not separate them.
The favourable one: lesson planning is genuinely time-consuming, teachers are using AI to draft and then applying judgement, and the output is better lessons in less time.
The uncomfortable one: the tools are being used where they are easiest to use rather than where the burden is heaviest, and administrative work, compliance reporting and data entry remain untouched because they involve systems an assistant cannot reach.
The 40% of hours spent on non-instructional tasks is the target. Lesson planning is not in that 40%.
What quality reporting adds
Teachers report quality improvements alongside time savings, and the pattern is informative.
Improvement is reported by 74% for administrative work, 64% for materials adapted to student needs, 61% for insights from student learning data, and 57% for grading and feedback, with 16% or fewer reporting reduced quality on any task.
The ordering is the interesting part. Reported quality gain is highest where the task is most routine and lowest where it requires judgement.
That is what one would predict if the tool is good at generation and weaker at evaluation, and it is consistent with the platform finding: the heaviest usage is in planning, and the lowest reported quality gain is in grading, which is where the previous article found agreement figures at their weakest.
The figures, sorted by how they were obtained
Separating logged from reported resolves most of the apparent disagreement.
| Figure | Source type | Reliability |
|---|---|---|
| 68% weekly use, from 29% | Survey, n=4,200 | Corroborated by logs |
| 1,438 to 4,644 teachers, 412,304 messages | Platform logs | Direct observation |
| Lesson planning 60 to 99% of usage | Platform logs | Direct observation |
| 5.9 hours saved weekly | Teacher recall, half-hour increments | Unverified |
| 10 to 14 hours saved weekly | Teacher recall | Unverified |
| 62% report partial offset | Teacher recall | Unverified, and directionally against interest |
Read the middle column and the article writes itself.
Everything observed directly concerns adoption and usage pattern. Everything concerning benefit is recall.
And the last row is the one worth weighting most heavily despite being unverified, because it runs against the respondent's interest. A teacher reporting that a tool saved them less than it appears to has no incentive to say so, which is the standard reason to trust a disclosure in this corpus.
The row that would resolve the subject does not exist. Non-instructional hours, measured before and after adoption, on the same teachers. That is a timesheet study, it is expensive, it is intrusive, and nobody has run it.
Which is measurement concentration making its second forward prediction in this territory. Platform logs are free and accrue automatically, so usage patterns are known precisely. Hours reclaimed require a separate instrument and a reason to build one, so they are surveyed rather than measured.
What the corpus can and cannot say here
Territory 12 has now produced three subjects where a second instrument existed and one where it does not, and the difference is instructive.
In tutoring, the second instrument was an unassisted post-test, which reversed the sign of the effect.
In detection, it was a demographic breakdown, which showed the errors sorting on prior attainment.
In grading, it was human-human agreement, which reframed what an 85% figure means.
Here there is nothing. No unassisted condition, no logged outcome, no independent baseline.
Which means this article can establish what teachers do and cannot establish what it is worth, and saying so is more useful than assembling the survey figures into a conclusion they cannot support.
The corpus should be explicit about the asymmetry that creates. Subjects with a checking instrument produce findings that look critical, because a second measurement usually qualifies the first. Subjects without one produce either credulous coverage or nothing.
And this corpus has been writing the first kind almost exclusively, which is a selection effect in its own output: the articles are about domains where somebody built the checking instrument, and the domains where nobody did are underrepresented here for the same reason they are underrepresented everywhere.
Three things this establishes
Adoption is real, corroborated and fast. 29% to 68% weekly in twelve months across a 4,200-teacher sample, with logged platform growth of 61% in a quarter. Survey and telemetry agree, which is rare in this corpus.
Time savings are self-reported with no independent instrument. 2.9 to 14 hours across sources, estimated by recall in half-hour increments, in a domain where the corpus has documented self-report running favourable three times and cannot check it here.
And the usage is not where the workload case points. Lesson planning at 60 to 99% of platform activity, with grading and data analysis least used, while the 40% of hours consumed by non-instructional tasks is the burden the argument rests on.
What it does not establish
That the time savings are false. 62% reporting a partial offset means a majority also report a real saving, and the direction is consistent across instruments.
That lesson planning use is misdirected. Planning is genuinely burdensome, and a teacher choosing to reduce it is making a judgement about their own week that no survey overrides.
That AI worsens burnout. No evidence here suggests that, and the strongest burnout predictor is the one the tools address.
And nothing about student outcomes. Every figure in this article is about teacher time and teacher perception. Whether students learn more is a separate question this literature does not ask.
What is unresolved
What the net saving is. The 62% offset is reported as a proportion of teachers, not as a quantity of time, so the headline figure cannot be adjusted.
Whether the hours are reinvested or reclaimed. Qualitative responses describe both more nuanced feedback and getting home earlier, which are different outcomes with different implications for retention.
Whether administrative burden actually falls. The 40% figure is the target and lesson planning is not in it, and no study measures non-instructional hours before and after.
And whether any of it affects retention. Attrition is the outcome the case is made on, early-career departures doubled between 2010 and 2024, and no study links AI adoption to a retention figure.
The counter-argument
Demanding telemetry in education imports a standard the field cannot meet. Teachers do not work in an instrumented environment the way developers commit to repositories, so criticising the evidence for being self-reported asks for a measurement that would require surveillance most teachers would rightly refuse. The survey data is the best obtainable.
The lesson planning finding may be the point rather than the problem. If planning is the most time-consuming professional task, teachers concentrating there are optimising correctly, and treating administrative work as the only legitimate target imposes an outsider's view of which hours are worth reclaiming.
The 62% offset figure is being over-read. Partially offset is not mostly offset, the same respondents reported a net saving, and this article promotes a caveat to a headline on the basis of a proportion whose magnitude is unknown.
And the burnout ordering argument cuts against the article's own framing. If administrative burden is the strongest predictor and AI reduces it, then a modest, partially offset, self-reported saving on the top predictor may matter more than a larger saving elsewhere, which the article treats as a limitation rather than as the case for the intervention.
The short version
Adoption is fast and well evidenced. RAND's January 2026 survey of 4,200 K-12 teachers found 68% using AI weekly against 29% a year earlier, and logged platform data across 12 districts shows 1,438 teachers in September rising to 4,644 by March, exchanging 412,304 messages.
Time savings are self-reported and vary fivefold. 2.9 hours a week for monthly users and 5.9 for weekly users in one survey, 10 to 14 hours in another, estimated by teachers in half-hour increments task by task. No education equivalent of repository telemetry exists to check them, in a corpus that has documented self-report running favourable three times where a second instrument was available.
RAND's own caveat is the line that should lead. 62% of teachers said the saving was partially offset by time spent reviewing outputs, which is the third domain showing this offset after clinical documentation and code review. In teaching the reviewer is the same person who saved the time, so both figures describe one week.
And the platform logs undercut the framing. Lesson planning is 60 to 99% of usage in every district, with grading and student data analysis least used, while the burden the case rests on is the 40% of hours spent on non-instructional tasks, which lesson planning is not part of.
The burnout ordering is the honest summary. Administrative load is the strongest predictor of attrition, which supports the intervention. Role ambiguity and lack of autonomy are second and third, and no tool touches either.
Common questions
How many teachers are using AI? Most, and adoption roughly doubled in a year. A RAND survey of 4,200 K-12 teachers across the United States in January 2026 found 68% using AI tools at least weekly in their professional practice, up from 29% in January 2025. A Gallup and Walton Family Foundation survey found 60% using AI over a school year with 30% weekly. Logged platform data across 12 districts corroborates the trajectory, rising from 1,438 teachers in September to 4,644 by 31 March 2026, a 61% increase.
How much time does it save? Reported savings range fivefold depending on source. Gallup and Walton report 5.9 hours a week for weekly users and 2.9 hours for monthly users, framed as about six weeks a year. RAND reports 10 to 14 hours a week for teachers using AI across multiple categories. Every figure is a teacher's estimate, with one methodology note specifying that savings were estimated by the half-hour on a task-by-task basis and summed across non-overlapping tasks.
Why does self-reporting matter here? Because this corpus has documented three cases where self-report ran systematically more favourable than measurement, and in each of those a second instrument existed to check it. Developers estimated 20% faster and were measured 19% slower; executives reported 97% benefiting against 29% organisational return; developers reported improved code quality while repository telemetry showed refactoring collapsing. In education there is no equivalent of repository telemetry, so the pattern can be flagged and not resolved.
What is the offset caveat? RAND's researchers noted that 62% of teachers reported the time saved was partially offset by time spent reviewing AI outputs. That is the third domain showing this pattern, after ambient clinical scribes where mandatory physician review partially offsets capture savings, and software where senior engineers report 20 to 35% more code review time. Education differs in who absorbs it: the reviewer is the same teacher who saved the time, so the reported saving and the reported offset describe the same person's week.
What do teachers actually use it for? Lesson planning, overwhelmingly. Logged platform data shows lesson planning features made up 60 to 99% of usage in every participating district, with frequent users concentrating even more heavily. Survey data agrees on ranking: research and content gathering at 44%, creating lesson plans at 38%, summarising at 38%, and generating classroom materials at 37%, with grading, one-on-one instruction and student data analysis least common.
Why does that complicate the workload case? Because the case rests on administrative burden. The OECD's survey across 48 countries found teachers spending 40% of working hours on tasks other than direct instruction, and burnout research ranks administrative load as the strongest predictor of attrition. Lesson planning is not part of that 40%: it is the creative and professional core of the job. Teachers may be using the tools where they are easiest to use rather than where the burden is heaviest, since compliance reporting and data entry involve systems an assistant cannot reach.
Can AI address teacher burnout? Partially, and the ceiling is structural. Burnout predictors rank administrative workload first, which AI does address, but role ambiguity and conflicting demands second and lack of autonomy third, neither of which a tool can touch. AI cannot clarify contradictory expectations from administrators, standardise inconsistent evaluation criteria, or return decision-making authority to a teacher who has none. As one analysis puts it, making unfulfilling work more efficient is not the same as making it meaningful.
What is the strongest objection to this article? That demanding telemetry imports a standard education cannot meet. Teachers do not work in an instrumented environment the way developers commit to repositories, so criticising the evidence for being self-reported effectively asks for surveillance most teachers would rightly refuse, and survey data is the best obtainable. A second objection is that the lesson planning finding may be the point rather than the problem: if planning is the most time-consuming professional task, teachers concentrating there are optimising correctly, and treating administrative work as the only legitimate target imposes an outsider's judgement about which hours are worth reclaiming.
Sources
Primary documents only. Where a claim rests on a single report, the entry says so.
- Generative AI in K-12 Classrooms: A Midyear Implementation Report arXiv:2605.16277 The logged platform data: 12 districts, 1,438 teachers in September rising to 4,644 by 31 March 2026, over 412,304 messages, 12 sessions and 3.7 hours per active user, and lesson planning at 60 to 99% of usage in every district.
- Three in 10 Teachers Use AI Weekly, Saving Six Weeks a Year Gallup and Walton Family Foundation The survey figures of 5.9 hours saved weekly for weekly users and 2.9 for monthly users, the task-by-task half-hour estimation method, and reported quality improvement from 57% for grading to 74% for administrative work.
- AI in Education 2026: The $32 Billion Market and What Teachers Actually Think AI Magicx The RAND January 2026 survey of 4,200 K-12 teachers: 68% weekly use against 29% a year earlier, 10 to 14 hours saved across multiple categories, and the caveat that 62% reported the saving was partially offset by time spent reviewing outputs.
- Reducing Teacher Burnout With AI: What the Research Says OpenEducat The OECD TALIS finding of 40% of working hours on non-instructional tasks across 48 countries, the burnout predictor ranking with administrative load first and autonomy third, and the observation that making unfulfilling work more efficient is not the same as making it meaningful.
Further reading
The primary literature behind the claims above, drawn from the concept entries this post links to, so a claim carries the same source here as it does there.
- METR (2025), randomised trial of experienced developers on their own repositories — self-estimated 20% speedup against a measured 19% slowdown. :: https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/ Self-Report Gap
- Wong et al. (2021), External Validation of a Widely Implemented Proprietary Sepsis Prediction Model — reported performance against independently measured performance on the population that mattered. :: https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2781307 Self-Report Gap
- Raji et al. (2021), AI and the Everything in the Whole Wide World Benchmark — how the availability of a benchmark shapes what a field concludes it has measured. :: https://arxiv.org/abs/2111.15366 Measurement Concentration
Related articles
- 72 seconds, or 30 minutes, and both are trialsTerritory 11 opens on the first subject this corpus has examined where the evidence is genuinely good. Registered trials, CONSORT-AI reporting, peer review, and effect sizes that still differ by a factor of twenty-five.
- Refactoring fell from 25% to 3.8%Survey evidence says AI improves code quality. Repository telemetry says the opposite. They are measuring different things, and the gap between them is where the productivity went.
- Clinicians override 49% to 96% of alertsFive years after the sepsis model was externally validated and found wanting, the sector-level evidence on clinical prediction alerts is process markers and no high-quality mortality signal.
- The same PDF says 83% and nobody quotes itTerritory 10 opens on enterprise deployment. The most-quoted statistic in the field is real, measures something much narrower than its use, and is contradicted inside its own source document.