Home/Blog/Evaluation & evidence/One trial, a waitlist control, and a letter

One trial, a waitlist control, and a letter

The best evidence for AI mental health support is a single randomised trial of a purpose-built clinical tool. Its own journal published three methodological objections, and almost nobody uses the thing that was tested.

TL;DR. The field's strongest evidence is one trial: Therabot, published in NEJM AI on 27 March 2025, randomising 210 adults to a four-week intervention or a waitlist control, reporting roughly 51% reduction in depression symptoms, 31% in anxiety and 19% in eating-disorder concerns, with therapeutic alliance rated comparable to outpatient psychotherapy. The same journal published a letter identifying three methodological limitations: the waitlist control, the absence of independent evaluation, and the use of a measure developed for human therapeutic relationships. And the tool tested is not the tool people use. Therabot was fine-tuned by clinicians over years. Independent testing found popular general-purpose models responding inappropriately to mental health symptoms at least 20% of the time, and a review of 160 chatbot studies found only 16% of LLM studies had undergone clinical efficacy testing.

---

Status: one good trial, formally contested, in a field where the evidence and the usage describe different products. Primary sources are the NEJM AI trial and the NEJM AI letter responding to it, alongside published safety research and a systematic review. This article reports research findings. It is not clinical guidance and makes no recommendation about any tool.

---

What the trial did

Heinz and colleagues conducted a national randomised controlled trial of 210 adults, published in NEJM AI on 27 March 2025.

Participants had clinically significant symptoms of major depressive disorder, generalised anxiety disorder, or were at clinically high risk for feeding and eating disorders, and were stratified into those three groups.

106 were assigned to a four-week Therabot intervention and 104 to a waitlist control, which received no app access during the study and gained it afterwards. Symptoms were assessed at baseline, at four weeks and at eight weeks.

Reported reductions against the control were roughly 51% for depression, 31% for anxiety and 19% for eating-disorder concerns, sustained at follow-up, with participants engaging for around six hours on average and rating their therapeutic alliance as comparable to outpatient psychotherapy.

Therabot is not a general chatbot. It was fine-tuned on cognitive behavioural therapy and psychotherapy practice by a clinical team over several years, and the researchers stated that clinician supervision is essential.

As a first randomised trial in this area, it is a substantial piece of work, and the corpus treats it as the strongest evidence available rather than as a target.

The letter in the same journal

**NEJM AI published a formal response identifying three methodological limitations that, in its authors' words, undermine confidence in the conclusions.**

The waitlist control. This is the substantive one. A waitlist group receives nothing, so the comparison captures expectancy, attention, engagement and natural symptom fluctuation alongside any treatment effect. Waitlist controls are known to produce larger effect sizes than active controls in psychotherapy research, which is why the design is generally regarded as establishing that something happened rather than that the specific intervention worked.

The absence of independent evaluation. The tool was assessed by the team that built it, which is the external validation problem this corpus has documented across several territories, most directly in the sepsis case.

And the misapplication of a measure developed for human therapeutic relationships. The therapeutic alliance instrument was constructed and validated to measure a relationship between two people. Applying it to a person and a chatbot produces a number, and whether that number means what it means in its original setting is precisely the construct validity question.

None of these is a claim that the trial is worthless. They are specific, published, and from a source with no commercial position.

What an active control would look like

A separate pilot randomised trial from Hong Kong shows the design the critique asks for.

124 participants were randomised one-to-one between an AI chatbot and a conventional nurse hotline, with 62 and 41 respectively completing pre- and post-questionnaires using GAD-7 and PHQ-9.

That comparison is against something rather than against nothing. It reports the chatbot showing potential in alleviating short-term anxiety and depression relative to the hotline, and its authors state plainly that more extensive randomised studies are needed.

It is a pilot, it is small, and its completion rates differ substantially between arms, which limits what it establishes.

But it is the right shape, and the gap between one good trial with a waitlist and one small pilot with an active control is the entire evidence base for whether these tools work better than an alternative.

The gap between what was tested and what is used

This is where the subject diverges from the rest of Territory 11.

Therabot was purpose-built, clinician-designed, fine-tuned over years, and studied under supervision.

Most people encountering "AI therapy" are using general-purpose chatbots, which were not designed for this, not tested for it, and in most cases positioned as wellness products outside any regulatory review.

Independent testing found popular models responding inappropriately to mental health symptoms at least 20% of the time.

A systematic review of 160 chatbot studies found LLM-based tools jumping to 45% of new studies in 2024, with only 16% of those studies having undergone clinical efficacy testing.

And professional assessment is sceptical: 94% of psychologists surveyed reported that chatbots cannot treat conditions with appropriate nuance.

So the evidence describes one product and the usage describes another, which is the same structure the shadow AI article found in enterprises: the sanctioned thing is studied and the actual thing is not.

The difference is the stakes. An unapproved productivity tool exposes data. A general-purpose chatbot responding to somebody in distress is a different category of failure, and the documented harms include cases involving minors, litigation, and settlements reached in January 2026.

This article does not describe those failure modes in detail, because doing so would be more useful to somebody constructing one than to somebody avoiding one.

Why the demand exists

Worth stating, because dismissing these tools without it misses why the question is urgent.

137 million Americans, roughly 40% of the population, live in a Mental Health Professional Shortage Area, according to federal designation as of December 2025.

That is the condition any assessment has to be made against. The comparator for a large share of potential users is not a therapist. It is nothing.

Which cuts both ways and is often used only one way. It is the strongest argument for developing these tools and the strongest reason the evidence bar matters, because a population with no alternative has no capacity to absorb a product that makes things worse.

And it explains the regulatory pattern. Illinois, Nevada and Utah became the first states to restrict or ban AI delivery of therapy in 2025, with six more advancing bills by early 2026. Those are restrictions on a product category in a market defined by scarcity, which is a hard position for a legislature and an honest reading of the evidence available to it.

Why the control group is the whole argument

The waitlist objection sounds procedural and is not, so it is worth setting out what it means for a 51% figure.

A waitlist group receives nothing and knows it. Over four weeks, several things happen to them that have nothing to do with any treatment.

Symptoms fluctuate. People typically enrol in mental health trials when they feel worse than usual, and symptom severity tends to drift back toward a personal average regardless of intervention. This alone produces improvement in untreated groups.

Expectancy does not apply to them. The intervention group knows it is receiving something intended to help. Expectation of benefit produces measurable symptom change in psychotherapy research, which is why active controls exist.

And attention is absent. Six hours of structured engagement with anything that responds is an intervention in itself, separate from the content of what it says.

A waitlist comparison therefore measures the sum of all four: symptom regression, expectancy, attention, and whatever the tool specifically does. Only the fourth is the product.

Which is why psychotherapy research treats waitlist-controlled effect sizes as systematically larger than active-controlled ones, and why a finding of this size against a waitlist is a reason to run the next trial rather than a measure of the intervention.

None of that makes 51% a wrong number. It makes it a number about a comparison, and the comparison was against nothing.

The practical consequence is specific. A prospective user deciding between a chatbot and a therapist, or between a chatbot and a support group, has no evidence bearing on either choice, because the trial did not make either comparison.

What this territory has now found four times

Territory 11 was chosen because its evidence is good, and every article has found a different failure that good evidence does not fix.

Ambient scribes: registered trials, and the effect varies twenty-five-fold by setting. Measurement moved the question rather than closing it.

The mammography trial: a hundred thousand participants and three Lancet papers, and the public claim outran the finding through a chain nobody controls. Better trials fix what is known and not what is repeated.

Drug discovery: every study registered and peer-reviewed, and the field-level rate disputed by twenty-eight points. Rigour is unit-scoped and nobody owns the sum.

LLM diagnosis: randomised trials, and the benchmark inherited from medical education tests the half of the task humans find hard. The evaluation was engineered before anyone thought to evaluate a machine on it.

And here: one trial, and its control group was nothing.

Five failures, none of which is a shortage of rigour. They are a setting effect, a transmission chain, a level mismatch, an inherited task boundary and a comparator choice.

Which suggests the corpus's recurring recommendation needs qualifying. This corpus has argued across five territories that better measurement would settle contested questions. Territory 11 is the test case, and what better measurement produced was better-specified uncertainty.

That is genuine progress and it is not what the recommendation promised, and saying so is more useful than repeating the recommendation.

Three things this establishes

The best evidence in the field is one trial with a design its own journal formally questioned. A waitlist control, developer evaluation, and a borrowed instrument are three specific objections published in NEJM AI, and none has been answered by a subsequent trial.

The evidence and the usage describe different products. A clinician-built tool studied under supervision is not what most people encounter, and the general-purpose models most people do encounter respond inappropriately at least a fifth of the time in independent testing and are mostly untested for clinical efficacy.

And the comparator is often nothing. With 40% of a population in a designated shortage area, an argument that a tool is worse than a therapist does not settle whether it is worse than the alternative available, which is the question a user faces and not the one the trials asked.

What it does not establish

That Therabot does not work. The trial is real, the effect sizes are large, and the objections concern what the design can support rather than whether anything happened.

That general-purpose chatbots always fail. Twenty percent inappropriate responses means eighty percent were not, and the testing measured specific scenarios rather than typical use.

That regulation is correct or incorrect. This corpus does not take positions on contested political questions, and state restrictions on AI therapy are one.

And nothing about any individual's care. No result here bears on what any person should do, and this article is not guidance.

What is unresolved

Whether the effect survives an active control. The single most important missing study, straightforward to design, and not yet run at scale.

Whether therapeutic alliance means anything here. The measure was built for human relationships and its application to chatbots is contested in the literature that uses it.

What general-purpose model performance actually is. Independent testing covers scenarios rather than populations, and no study measures outcomes among people using consumer chatbots for support.

And whether supervision is achievable at scale. The Therabot researchers stated clinician supervision is essential, and the deployment model most people encounter has none.

The counter-argument

A waitlist control is standard and often necessary. Withholding an active comparator raises its own ethical questions in a first trial of a novel intervention, and criticising a study for not running the harder design ignores why the easier one is conventional at this stage. The letter identifies real limitations and does not establish that the trial should have been done differently.

The 20% figure is doing heavy work here. It comes from scenario-based testing rather than from observed use, the scenarios were selected to probe failure, and a rate measured on adversarially chosen prompts is not a rate users experience. This article uses it as though it described typical performance.

Comparing the trial tool to consumer chatbots may be unfair to both. Therabot's evidence was never offered as evidence about general models, and criticising the field for a gap between them holds researchers responsible for products they did not build and explicitly distinguished themselves from.

And the shortage argument can justify too much. "Better than nothing" is the reasoning behind most poorly evidenced interventions in medicine, and the history of that argument is not encouraging. A population with no alternative is the population least able to bear a harm, which this article states and then partly sets aside.

The short version

The field's strongest evidence is one trial. Therabot, NEJM AI, 27 March 2025: 210 adults, 106 to a four-week intervention and 104 to a waitlist control, reporting roughly 51% reduction in depression symptoms, 31% in anxiety, 19% in eating-disorder concerns, sustained at eight weeks, with alliance rated comparable to outpatient psychotherapy.

The same journal published a letter identifying three limitations: the waitlist control, which captures expectancy and natural fluctuation alongside treatment effect; the absence of independent evaluation; and a therapeutic alliance measure built for relationships between people.

A separate pilot in Hong Kong shows the design the critique asks for, randomising 124 participants between a chatbot and a nurse hotline, which is a comparison against something rather than nothing.

And the tool tested is not the tool used. Therabot was clinician-built over years with stated supervision requirements. Independent testing found popular general-purpose models responding inappropriately to mental health symptoms at least 20% of the time, a review of 160 chatbot studies found only 16% of LLM studies clinically tested, and 94% of surveyed psychologists reported chatbots cannot treat with appropriate nuance.

The demand is not in doubt. 137 million Americans, about 40%, live in a designated Mental Health Professional Shortage Area, which makes the comparator for many people nothing at all, and makes the evidence bar more important rather than less.

---

If you are struggling with your mental health, support is available, and a conversation with a doctor or a local service is a better starting point than anything described here.

Common questions

What did the Therabot trial find? A national randomised controlled trial of 210 adults, published in NEJM AI on 27 March 2025, assigned 106 participants to a four-week intervention with the Therabot app and 104 to a waitlist control. Participants had clinically significant symptoms of major depressive disorder, generalised anxiety disorder, or were at clinically high risk for feeding and eating disorders. Reported reductions against control were roughly 51% for depression, 31% for anxiety and 19% for eating-disorder concerns, sustained at eight-week follow-up, with around six hours of average engagement.

What were the objections to it? NEJM AI published a letter identifying three methodological limitations. The waitlist control, which means the comparison captures expectancy, attention and natural symptom fluctuation alongside any treatment effect, and which is known to produce larger effect sizes than active controls in psychotherapy research. The absence of independent evaluation, since the tool was assessed by the team that built it. And the use of a therapeutic alliance measure developed and validated for relationships between people, applied to a person and a chatbot.

Is Therabot the same as a general chatbot? No, and the distinction matters more than anything else in this subject. Therabot was fine-tuned on cognitive behavioural therapy and psychotherapy practice by a clinical team over several years, and its researchers stated that clinician supervision is essential. Most people encountering AI mental health support are using general-purpose chatbots that were not designed for it, not tested for it, and are mostly positioned as wellness products outside regulatory review.

How do general-purpose models perform? Poorly in independent testing, though the figure should be read carefully. Popular models responded inappropriately to mental health symptoms at least 20% of the time in published scenario-based research. A systematic review of 160 chatbot studies found LLM-based tools rising to 45% of new studies in 2024 with only 16% of those studies having undergone clinical efficacy testing, and 94% of surveyed psychologists reported that chatbots cannot treat conditions with appropriate nuance.

Why does the shortage matter to the assessment? Because it sets the comparator. Roughly 137 million Americans, about 40% of the population, live in a designated Mental Health Professional Shortage Area as of December 2025. For a large share of potential users the alternative to a chatbot is not a therapist but nothing at all. That is simultaneously the strongest argument for developing these tools and the strongest reason the evidence bar matters, because a population with no alternative has the least capacity to absorb a product that makes things worse.

What is the regulatory position? Illinois, Nevada and Utah became the first states to restrict or ban AI delivery of therapy in 2025, with six more advancing bills by early 2026. This corpus does not take positions on contested political questions, and restrictions on AI therapy are one. What can be said is that these are restrictions on a product category in a market defined by scarcity, which is a genuinely difficult position for a legislature.

What would settle the effectiveness question? A trial with an active control. The Therabot result compares against a waitlist, so it establishes that something happened rather than that the specific intervention was responsible. A pilot from Hong Kong randomising 124 participants between a chatbot and a nurse hotline shows the right design at small scale. A larger trial comparing against an existing service, or against a non-specific supportive intervention, is the missing study and is straightforward to design.

What is the strongest objection to this article? That the 20% inappropriate-response figure is doing heavy work. It comes from scenario-based testing where prompts were selected to probe failure, which is not a rate users experience in typical conversation, and this article uses it as though it described general performance. A second objection is that comparing a clinician-built research tool to consumer chatbots holds researchers responsible for products they did not build and explicitly distinguished their work from.

Sources

Primary documents only. Where a claim rests on a single report, the entry says so.

  1. Randomized Trial of a Generative AI Chatbot for Mental Health Treatment Heinz et al., NEJM AI, 27 March 2025 The trial: 210 adults, 106 to a four-week Therabot intervention and 104 to a waitlist control, stratified by depression, anxiety and eating disorder risk, with outcomes at four and eight weeks.
  2. A Letter about Randomized Trial of a Generative AI Chatbot for Mental Health Treatment NEJM AI The published response identifying three methodological limitations: the waitlist control, the absence of independent evaluation, and the misapplication of a measure developed for human therapeutic relationships.
  3. Comparison of an AI Chatbot With a Nurse Hotline in Reducing Anxiety and Depression Levels Chen et al., pilot randomised controlled trial, 2025 The active-control design the critique asks for: 124 participants randomised between an AI chatbot and a conventional nurse hotline, using GAD-7 and PHQ-9, with the authors stating more extensive trials are needed.
  4. AI Therapy Statistics 2026: Usage, Effectiveness and Safety Psychology.com The surrounding figures: independent testing finding popular models responding inappropriately to mental health symptoms at least 20% of the time, only 16% of LLM chatbot studies clinically tested, and 137 million Americans in designated shortage areas.

Further reading

The primary literature behind the claims above, drawn from the concept entries this post links to, so a claim carries the same source here as it does there.

  • Wong et al. (2021), External Validation of a Widely Implemented Proprietary Sepsis Prediction Model — the canonical case, and the source of the 7% figure. :: https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2781307 External Validation
  • Raji et al. (2021), AI and the Everything in the Whole Wide World Benchmark — why a developer-selected evaluation cannot license a general claim. :: https://arxiv.org/abs/2111.15366 External Validation
  • Bowman & Dahl (2021), What Will it Take to Fix Benchmarking in Natural Language Understanding? — what a benchmark must satisfy to support inference. :: https://arxiv.org/abs/2104.02145 Construct Validity

Learn the concepts

← All posts