Home/Blog/Human in the loop is weaker than it sounds

Human in the loop is weaker than it sounds

When the AI was wrong, experienced radiologists went from 82% accurate to 45.5%. A review step is only a control if the reviewer sometimes disagrees, and almost nobody measures how often they do.

A 2023 study put twenty-seven radiologists in front of fifty mammograms alongside AI suggestions. When the AI was correct, everything worked as advertised.

When the AI was wrong, inexperienced radiologists' accuracy fell from around 80% to under 20%. The experienced ones, averaging more than fifteen years in the specialty, fell from 82% to 45.5%.

Expertise did not protect them. It halved them.

"A human makes the final decision" is the sentence that unlocks most AI deployments in regulated settings, and the evidence that it works is considerably weaker than the confidence placed in it. A review step is a control only if the reviewer sometimes disagrees. The rate at which they do is measurable, it is almost never measured, and where it has been measured it is low enough to make the control largely decorative.

What the evidence actually shows

Radiology has been studied more than any other domain because the setup is clean: a decision with a ground truth, a specialist making it, and an AI making a recommendation. The findings are consistent and they are not encouraging.

A systematic review of studies published between 2016 and 2026, covering mammography, chest radiography and MRI, found radiologists following incorrect AI recommendations at high rates across every included study, with a pooled odds ratio of 4.89. Following the AI when the AI was wrong was roughly five times more likely than not.

A controlled experiment published in 2026 measured it directly by perturbing 30% of an AI's recommendations by one category, enough to be plausible and not enough to be obviously wrong, without telling the radiologists. Under standard AI assistance, automation bias occurred in 36.1% of the manipulated cases. Roughly a third of the time, a wrong recommendation carried the reader with it.

The same study measured something more uncomfortable. When the AI was revealed after the radiologist's own initial read, 33.9% revised a correct first impression toward the wrong AI recommendation. Deciding first does not protect you. It just changes which bias applies.

Outside imaging, a study of UK general practitioners found clinicians changing prescriptions in response to decision-support advice in about 22.5% of cases, and in 5.2% of all cases switching from a correct prescription to an incorrect one after receiving erroneous advice. In computational pathology, over 30% of participants reversed correct initial diagnoses when shown incorrect AI output.

None of these are studies of careless people. They are studies of trained specialists doing the task they trained for, with a review mechanism in place, failing in a consistent direction.

Three mechanisms, and they are not the same thing

Treating this as one phenomenon leads to one intervention, which is why interventions usually fail. There are at least three, with different causes and different fixes.

Automation bias is deference to a source perceived as authoritative. It is documented across aviation, radiology, criminal justice and hiring, long before the current wave of systems. The mechanism is that an automated recommendation carries an implied warrant, and the reviewer's own uncertain judgement is weighed against something that appears certain.

Decision fatigue is volumetric and it is empirically distinct. A reviewer processing hundreds of items per shift anchors on the first few, applies decreasing scrutiny as the queue lengthens, and eventually treats approval as the default because approval is the path of least resistance. This is caused by volume, not by perceived authority, and it will occur even if the reviewer holds the system in contempt.

Anchoring operates on the reviewer's own prior judgement. Once an initial read exists, a conflicting AI recommendation does not prompt a fresh analysis; it prompts a revision, and revisions run toward the more recent input. This is why "have the human decide first" is a weaker safeguard than it appears, as the 33.9% revision figure shows.

The practical consequence: an intervention aimed at one does nothing for the others. Reducing queue volume addresses fatigue and not deference. Hiding the AI's recommendation until after the human decides addresses deference and creates an anchoring problem instead.

Omission and commission

The failures also split by type, and only one of them is visible.

Commission errors are following the AI against contradicting evidence. The reviewer had reason to disagree and did not. These are at least detectable after the fact, because the contradicting evidence was in the record.

Omission errors are failing to notice something the AI missed. The reviewer approved an output whose defect was an absence, and absences do not appear in a review log. This is the more common failure in fast-paced settings and the one no audit will surface, because there is nothing to audit.

An organisation reviewing its incident history will find commission errors and conclude that its reviewers occasionally follow bad advice. It will not find the omission errors, and it will therefore underestimate the problem by an unknown margin.

Does explaining the AI help? Contested.

The intuitive fix is to show the reviewer why the system reached its conclusion, on the theory that visible reasoning invites scrutiny.

The 2026 radiology study supports this substantially. Adding saliency heatmaps alongside recommendations cut automation bias from 36.1% to 17.8% and anchoring bias from 33.9% to 17.2%, both statistically significant, with adjusted odds ratios around 0.56 and 0.61. On the unmanipulated cases, accuracy improved from 86.2% unaided to 90.1% with AI plus explanation. That is a real result and it is the strongest evidence available for explanation as a mitigation.

Against it sits a persistent finding from the human-factors literature that explanations can increase trust without increasing accuracy, because a plausible-sounding rationale is itself persuasive. A confident explanation for a wrong answer may be worse than a bare wrong answer, since it supplies the reviewer with reasons to agree.

Both can be true. A saliency map showing where a model looked is a different artefact from a natural-language rationale explaining why it concluded something, and the former is much harder to fabricate convincingly. The distinction worth holding is between explanations that expose the computation and explanations that narrate it. The evidence for the first is decent. The evidence for the second is not, and language models produce the second by default.

The measurement almost nobody takes

Here is the diagnostic, and it is one number.

What proportion of items does your reviewer change?

If a reviewer approves 99.8% of what they see, one of two things is true. Either the system is right 99.8% of the time, which is a claim you can test independently and almost certainly cannot support. Or the review is not functioning, and you are recording a human decision that is not occurring.

The number to compare it against is the system's measured error rate on a held-out sample. If the model is wrong 4% of the time and your reviewers change 0.3% of outputs, the review is catching roughly one in thirteen of the errors it exists to catch, and the other twelve are being ratified by a person whose approval now appears in the record.

Two refinements make it sharper.

Seed known errors. Deliberately inject a small number of incorrect outputs into the review queue and measure how many are caught. This is the only method that measures the review rather than inferring it, and it is standard practice in other quality-assurance settings and nearly absent here.

Track it over time. Disagreement rates decay. A reviewer new to a system scrutinises; the same reviewer six months later has learned that the system is usually right, which is a rational update that also degrades the control. A control that erodes predictably needs to be measured continuously rather than validated once.

The arithmetic of a review layer

The claim "a human checks it" implies a specific improvement, and that improvement can be calculated. Doing so usually deflates it.

Suppose a model is correct 96% of the time, and a reviewer catches half the errors it makes. That is a generous assumption: the radiology evidence suggests something closer to a third, and only for errors the reviewer had reason to question.

Without review: 4 errors per 100 reach the outcome. With review at 50% catch rate: 2 errors per 100 reach the outcome, now carrying a human approval.

The system halved its errors. It also converted every surviving error from a machine failure into a human-approved decision, which is a different thing legally and organisationally even though the outcome is identical.

Now vary the catch rate, because that is the parameter nobody measures.

Reviewer catch rateErrors reaching outcomeEffective accuracy
90%0.4 per 10099.6%
50%2 per 10098%
30%2.8 per 10097.2%
10%3.6 per 10096.4%
0%4 per 10096%

The gap between a review layer working well and one that has decayed into ratification is the difference between 99.6% and 96%, which sounds small and is a nine-fold difference in error volume.

Two things follow. The catch rate is the whole value of the review, and it is the one quantity most deployments never establish. And a review layer at a low catch rate is not neutral: it costs the reviewer's time, it slows the workflow, and it supplies documentation that a human decided, which makes the remaining errors harder to attribute and easier to defend.

What actually keeps review meaningful

The interventions with support, stated with appropriate uncertainty since the evidence base is thinner than the problem warrants.

Cap the queue. Fatigue is volumetric, so the fix is volumetric. A reviewer handling forty items carefully is worth more than one handling four hundred by reflex, and the second arrangement produces better throughput numbers and worse outcomes.

Sample rather than review everything. Reviewing 10% of output attentively beats reviewing 100% of it by pattern-matching. Full review is frequently a compliance artefact rather than a control, and it converts an expensive specialist into a clicking mechanism.

Make disagreement cheap and agreement effortful. Most interfaces make approval one click and rejection a form. That gradient is a design decision and it produces the outcome it rewards. Requiring a brief reason for approval on high-stakes items inverts it, at a real cost in throughput.

Withhold the recommendation until the human commits. Imperfect, since it substitutes anchoring for deference, but the 33.9% revision rate is still lower than the 36.1% automation-bias rate, and the human's independent judgement is at least recorded before contamination.

Route by uncertainty rather than by rule. Sending everything to review dilutes attention across items that did not need it. Sending the model's least confident outputs concentrates it, provided the confidence estimate is calibrated, which is a separate problem and frequently unsolved.

Rotate and re-train. Deskilling is real: the scoping literature describes erosion of competence through cognitive offloading, as the practitioner shifts from doing the task to supervising it. A reviewer who no longer performs the underlying task unaided will eventually be unable to detect subtle error in it, and periodic unaided practice is the only known counter.

What aviation learned, and why it took thirty years

The strongest reason for cautious optimism is that another industry has already been through this, at higher stakes, and came out with procedures rather than platitudes.

Autopilot and flight management systems produced the same pattern in the 1980s and 1990s: crews deferring to automation against their own reading of the situation, monitoring performance degrading with time on task, and skills eroding as hand-flying became rare. Several fatal accidents were attributed to it. The response was not better automation. It was a set of institutional changes that treat the human layer as something requiring maintenance rather than something you install.

Four of them transfer.

Mandatory unaided practice. Pilots hand-fly on a schedule regardless of whether the automation is available, specifically to prevent the skill decay that makes supervision ineffective. The equivalent in an AI workflow is periodically doing the task without the model and comparing, which almost nobody does because it looks like waste.

Explicit mode awareness. A great many automation accidents involved a crew that did not know what the system was currently doing. The AI equivalent is a reviewer who does not know which version, which prompt, or which configuration produced the output in front of them, which is the normal situation rather than the exception.

Cross-checking as procedure, not attitude. Callouts and confirmations are scripted rather than left to vigilance, because vigilance is not reliable. The AI equivalent would be requiring a specific check on specific fields rather than asking for general scrutiny.

Non-punitive reporting. Crews report their own errors and near-misses without penalty, which is how the failure data exists at all. Most AI deployments have no equivalent, which is why organisations know their model's benchmark accuracy and not their reviewers' catch rate.

The uncomfortable part of the comparison is the timescale. Aviation took roughly three decades and a number of accidents to arrive at these, and it had a regulator, a shared incident database and a professional culture that treats procedure as identity. Most organisations deploying AI review have none of those.

The feedback loop nobody plans for

One consequence deserves separate mention because it compounds silently.

Approved outputs frequently become training data. If reviewers ratify a class of error, that error enters the next training cycle labelled as correct, and the model becomes more confident in exactly the behaviour the review failed to catch. The oversight mechanism has now amplified the defect it was installed to prevent.

This is a slow failure and it is invisible in any single cycle. The defence is keeping a held-out evaluation set that was never touched by the review pipeline, which is straightforward to state and organisationally difficult, because the reviewed data is the convenient data.

What is unresolved

Whether meaningful oversight is achievable at production volume. Every intervention above trades throughput for scrutiny. Whether there is a configuration that preserves both at commercial scale, or whether the honest answer is that high-volume oversight is a contradiction, is not settled. The literature documents the failure well and the successful counter-examples are thin.

Whether explanation helps or persuades. The radiology result is strong and specific to visual saliency. Whether it generalises to natural-language rationales, which are the dominant form in current systems and are optimised for plausibility, is untested and there are reasons to expect the opposite.

Whether the human is there for oversight or for liability. The uncomfortable possibility is that the review layer functions primarily to establish that a person was accountable, and that its detection performance is incidental to its purpose. If so, measuring detection would be beside the point, and the honest framing of many deployments would be different from the stated one. This is a claim about institutions rather than about psychology, and it is not testable in the same way.

The counter-argument

The radiology evidence may not transfer. These are perceptual judgements under time pressure with a single correct answer, which is a specific kind of task. Reviewing a drafted email, a code change or a classification decision may behave differently, and assuming the mammography numbers describe your workflow is an extrapolation.

Some review is better than none. A reviewer catching one error in thirteen still catches one in thirteen. The argument that imperfect oversight is decorative can slide into the argument that it is worthless, and that does not follow. The correct response to a weak control is usually to strengthen it rather than remove it.

Automation bias is a known problem with known mitigations. Aviation has spent decades on exactly this and developed procedures, cross-checks and training that meaningfully reduce it. The pessimistic reading treats the problem as novel when it is not, and the field it is borrowed from has answers worth importing.

And the base rate matters. If a system is right 99% of the time and a reviewer catches a third of the remaining errors, the combined system may outperform the unaided human by a wide margin even with badly degraded review. Measuring the review in isolation can produce a discouraging number about an arrangement that is working.

The short version

The claim that a human makes the final decision underwrites most AI deployment in regulated settings, and the evidence for it is weak. In controlled mammography studies, incorrect AI recommendations took inexperienced radiologists from around 80% accuracy to under 20%, and experienced radiologists with fifteen or more years from 82% to 45.5%. A systematic review across imaging modalities found a pooled odds ratio of 4.89 for following incorrect recommendations. In general practice, 5.2% of all prescribing cases involved switching from a correct to an incorrect prescription after erroneous decision-support advice.

Three distinct mechanisms are at work and they need different fixes. Automation bias is deference to perceived authority. Decision fatigue is volumetric and occurs regardless of how the reviewer regards the system. Anchoring operates on the reviewer's own prior judgement, which is why having the human decide first helps less than expected: 33.9% revised a correct initial read toward a wrong AI recommendation. Failures also split into commission errors, which the record captures, and omission errors, which by construction it cannot.

Explanation may help. Saliency heatmaps roughly halved both bias types in a controlled study, and accuracy on unmanipulated cases rose from 86.2% to 90.1%. Whether this transfers to natural-language rationales, which are optimised for plausibility, is untested and there are grounds for doubt.

The diagnostic is one number: what proportion of items does your reviewer change? Compare it against the system's measured error rate on held-out data. If the model errs 4% of the time and reviewers change 0.3% of outputs, the review is catching roughly one error in thirteen, and the remaining twelve now carry a human approval in the record. Seeding known errors into the queue is the only way to measure this rather than infer it, and almost nobody does it.

Common questions

Does human review actually catch AI errors? Less often than assumed. Controlled studies in radiology found specialists following incorrect AI recommendations roughly a third of the time, with a pooled odds ratio of 4.89 across a systematic review. In one mammography study, experienced radiologists' accuracy fell from 82% to 45.5% when the AI was wrong. Review catches some errors and the fraction is far below what "a human makes the final decision" implies.

What is automation bias? The tendency to defer to automated recommendations, weighting them above one's own judgement and above contradicting evidence. It is documented across aviation, radiology, criminal justice and hiring, and it predates current AI systems. It is distinct from decision fatigue, which is caused by volume rather than perceived authority, and from anchoring, which operates on the reviewer's own prior judgement.

Does having the human decide first prevent the problem? It helps and does not solve it. In a controlled study, when AI advice was revealed after the radiologist's initial read, 33.9% revised a correct first impression toward the wrong recommendation. That is lower than the 36.1% automation-bias rate when AI came first, so the ordering is a genuine improvement, but the human's judgement is still substantially revisable and the mechanism has shifted rather than disappeared.

How do I tell if my human review is working? Measure the proportion of items reviewers change, and compare it against the system's measured error rate on a held-out sample. If the model is wrong 4% of the time and reviewers change 0.3% of outputs, review is catching roughly one error in thirteen. To measure rather than infer, seed known incorrect outputs into the review queue and count how many are caught. Track the rate over time, since disagreement decays as reviewers learn the system is usually right.

Does showing the AI's reasoning help reviewers catch errors? Contested, and the answer may depend on the kind of explanation. A 2026 controlled study found saliency heatmaps roughly halving both automation bias and anchoring bias, with accuracy on unmanipulated cases improving from 86.2% to 90.1%. Against that, human-factors research finds explanations can raise trust without raising accuracy, since a plausible rationale is itself persuasive. Explanations that expose the computation appear to work better than explanations that narrate it, and language models produce the second kind by default.

What is the difference between omission and commission errors? Commission errors are following the AI despite contradicting evidence, which the record can capture because the evidence was there. Omission errors are failing to notice something the AI missed, which no audit surfaces because the defect is an absence. Reviewing your incident history will find commission errors and miss omission errors entirely, so it will underestimate the problem by an unknown amount.

Why does review quality degrade over time? Two reasons. Reviewers rationally update toward trusting a system that is usually right, which erodes scrutiny without any lapse in professionalism. And deskilling: shifting from performing a task to supervising it reduces practice at the underlying skill, so the reviewer becomes progressively less able to detect subtle error. Periodic unaided practice is the only known counter, and continuous measurement matters more than one-time validation.

What design changes make human review more effective? Cap queue volume, since fatigue is volumetric. Sample attentively rather than reviewing everything by reflex, because full review is frequently a compliance artefact. Invert the effort gradient, since most interfaces make approval one click and rejection a form, which rewards approval. Withhold the recommendation until the human commits, accepting that this substitutes anchoring for deference. Route by model uncertainty rather than reviewing uniformly. And keep an evaluation set the review pipeline never touches, because approved outputs frequently become training data and a ratified error will be reinforced.

Learn the concepts

← All posts