Comparator Choice
What a result was measured against, which is frequently unstated and often nothing, and which determines what the result can support.
When not to use it
- Where the comparator is stated, appropriate and matched to the reader's decision, which is the condition the concept exists to check for.
- As an objection to early-stage trials, where a no-treatment control is conventional and sometimes the only ethical option.
- Where no meaningful alternative exists, in which case a description is the honest output and should be labelled as one.
Reach for something else instead
- Active control — comparison against an existing option, which answers the question a user has.
- Stated base rate — publishing the corresponding rate for the incumbent, which costs one line and makes a figure interpretable.
- -
Read more on the blog
- One trial, a waitlist control, and a letterThe best evidence for AI mental health support is a single randomised trial of a purpose-built clinical tool. Its own journal published three methodological objections, and almost nobody uses the thing that was tested.
- AI support: 90% deflected, 40% actually resolvedA support system can report 90% deflection on a 40% resolution rate, because a customer who gets an unhelpful answer and gives up counts as deflected. The most-cited deployment reversed.
- The same PDF says 83% and nobody quotes itTerritory 10 opens on enterprise deployment. The most-quoted statistic in the field is real, measures something much narrower than its use, and is contradicted inside its own source document.
- Good evidence, and eight different failuresTerritory 11 closes. This was the field chosen because its evidence is strong, and every subject produced a failure that stronger evidence would not have prevented.
Further reading
- Heinz et al. (2025), Randomized Trial of a Generative AI Chatbot for Mental Health Treatment, and the published response identifying the waitlist control among three limitations.
- Wong et al. (2021), External Validation of a Widely Implemented Proprietary Sepsis Prediction Model — performance against the population that mattered rather than against overall discrimination.
Primary sources, listed so you can check the claims on this page rather than take them on trust.
Where people go wrong
- Comparing effect sizes from studies with different control arms as though they measured the same thing.
- Reading a figure with no comparator as an effect rather than as a description.
- Assuming the comparator used was the alternative the reader actually faces, which is frequently not the case.
At a glance
Where this sits
A starting point. Nothing needs to come before it.
Computed from the prerequisite graph, not assigned. How this works