Human-in-the-Loop
Putting a person at the decision point — the only reliable safeguard for agents, and it fails quietly when the person becomes a rubber stamp.
When not to use it
- At high volume. Ninety approvals a day is a clicking exercise, not a review.
- When the reviewer can't tell good from bad. You've added cost and a false sense of safety.
- On reversible, low-stakes actions. Save the attention for the decisions that need it.
- As the whole safety argument. "There's a human in the loop" should start the conversation, not end it.
Reach for something else instead
- Capability restriction — don't grant the action. More reliable than reviewing it.
- Post-hoc sampling — audit a fraction, if actions are reversible.
- Automated verification — a test is a better check than a tired person.
- Escalation on genuine uncertainty — if you can calibrate it, which is the hard part.
Read more on the blog
- Human in the loop is weaker than it soundsWhen the AI was wrong, experienced radiologists went from 82% accurate to 45.5%. A review step is only a control if the reviewer sometimes disagrees, and almost nobody measures how often they do.
- How to secure an LLM application: risks and defensesTraditional security rests on a boundary between instructions and data. A language model has no such boundary, because both arrive as one stream of text it cannot tell apart. That single fact determines every defense that works, and every one that does not.
- Will AI take my job? What the evidence showsThe frightening headline numbers and the reassuring ones are both real, because they measure different things. AI acts on tasks, not jobs, and a job is a bundle of tasks. That single distinction explains why the studies appear to contradict each other and what the evidence actually supports.
- Why AI agents fail: the seven failure modesGartner predicts over 40% of agentic AI projects will be canceled by 2027. The failures follow patterns, seven of them. The taxonomy: what breaks, why, which real incident proved it, and which control would have prevented it.
Further reading
- Parasuraman & Riley (1997), Humans and Automation: Use, Misuse, Disuse, Abuse — automation bias, from decades before anyone needed it for this.
- Bansal et al. (2021), Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance — human-AI teams often underperform the AI alone.
- Amershi et al. (2019), Guidelines for Human-AI Interaction — the practical design guidance, and it's specific.
Primary sources, listed so you can check the claims on this page rather than take them on trust.
Where people go wrong
- Reviewing everything, so nothing is reviewed. Rare and informative beats frequent and reflexive.
- Showing "Approve?" without the information needed to decide. You've asked for a reflex.
- Making rejection harder than approval. You've built a yes machine.
- Assuming a human improves the system. Bansal et al.: teams often underperform the AI alone.
- Treating high reliability as reassuring. It's what makes the reviewer stop looking.
At a glance
Often compared with
Where this sits
25 concepts come first. Understanding it opens up 1 more.
Computed from the prerequisite graph, not assigned. How this works