The human who approves everything
There’s a reviewer somewhere approving four hundred AI outputs a shift.
They approve almost all of them. They have to — the queue doesn’t stop, and nobody thanks them for being slow.
On the architecture diagram, that person is a control. In reality they’re a formality. The system has human-in-the-loop AI in exactly the way a fire door propped open has a fire door.
That’s the whole distinction. It isn’t a people problem — it’s a design problem, and you can measure whether yours works.
In one breath: Adding a reviewer creates a step, not a check. A check is a point where the chain can actually be stopped.
Three forces collapse it: volume, automation bias, and incentives that reward throughput. A 2006 review of clinical alerting found override rates of 49–96% — high-volume, low-precision review destroys itself.
So bound the volume, route by consequence rather than confidence, gate at chain boundaries, and make each escalation carry what the agent tried and why it’s unsure.
Then seed known-wrong outputs and measure the catch rate. Oversight you haven’t tested is a claim, not a control.
Oversight is a design property, not a staffing decision
“We’ll have a human review it” usually ends a risk conversation without changing a system.
A step is a point where the chain pauses and a person does something. A check is a step that can stop the chain — where “no” is a real, supported, consequence-free outcome for the person saying it.
Three questions separate them:
- Can the reviewer actually say no, and does the system handle that path?
- Do they have enough context to know when to say no?
- Is their workload low enough that saying no is possible?
Five setups below. Make the call before you read the reasoning — a guess you’ve committed to is one you remember being wrong about.
Interactive: five review setups to call as a step or a check, with the reasoning for each. Needs JavaScript.
Everything below is how to build the second kind.
The three positions, and what each commits you to
“Human in the loop” covers three genuinely different architectures. The difference is where the person sits relative to the action.
| Position | Who decides | Execution model | The failure it invites |
|---|---|---|---|
| In the loop | The human. AI recommends, nothing proceeds without input | Synchronous — the workflow pauses and waits | Rubber-stamping under volume; unbounded latency |
| On the loop | The AI. Humans monitor and hold veto | Asynchronous — dashboards and override surfaces | Nobody watching the dashboard; discovering failures late |
| Out of the loop | The AI, inside boundaries set at design time | Fully autonomous | Boundaries that were wrong, discovered in production |
Each position commits you to a different system, not a different policy.
- In the loop — the workflow stops dead and resumes where it left off. Someone goes to lunch and your system has to survive it.
- On the loop — a dashboard nobody opens fails the same way a reviewer who approves everything does.
- Out of the loop — right when a mistake is small and easy to undo. Claiming oversight you don’t have hides the problem somewhere harder to find.
- IN the loop
- ON the loop
- OUT of the loop
- Tagging support tickets by topic
- Sending a £4,200 refund
- A pricing model updating 40,000 listings overnight
- Deleting a customer's account
- Autocompleting a search box
- An agent posting status updates to a team channel
Why oversight collapses: the evidence
“People rubber-stamp AI” gets asserted constantly and evidenced almost never.
The override numbers
Clinical decision support has run this experiment at scale for twenty years: drug-safety alerts fire, a clinician responds, every response is logged.
In 2006 van der Sijs et al. pulled together seventeen studies of it. Clinicians overrode safety alerts in 49% to 96% of cases.
The part usually left out: the same review found overriding is often justified. The alerts fired too readily, on interactions that didn’t matter, to people who knew better.
So the lesson isn’t that clinicians stopped caring. It’s that high-volume, low-precision review destroys itself.
Interactive: two dials — items sent to review per shift, and the seconds a careful review needs — driving the share of flawed work that gets caught. Needs JavaScript.
Drag the volume up and the gate is still there. It just stops doing anything — the reviewer adapts correctly to a badly designed system, and every item becomes noise, including the one that mattered.
If your AI escalates everything, that’s the system you’re building.
Who actually rubber-stamps
Pause & recall. You’re staffing a review queue. Who is most likely to wave through a wrong AI output — someone who has never used AI, someone who uses it daily but isn’t technical, or an AI expert? Commit to an answer before you read on.
Reveal the answer
The one in the middle. Automation bias peaks at low-to-moderate AI familiarity — enough exposure to trust the output, not enough to know how it fails. Complete novices lean towards distrusting it. Experts trust it about as much as it deserves. The awkward part is that the middle profile is the obvious one to staff a review queue with.
Horowitz & Kahn tested 9,000 adults across nine countries. They published the method before running it, so the story couldn’t be fitted to the result.
The line rose, peaked, and came back down. The reviewer most likely to defer to a wrong AI output is the moderately AI-familiar one — comfortable with the tools, not deep in how they fail. That profile is the obvious one to put on a review queue.
The three forces, together
- Volume. Attention per item is total attention divided by item count. You control the denominator.
- Automation bias. Over-weighting machine output against contradicting evidence, documented across aviation, radiology and hiring.
- Incentives. If throughput is measured and catches aren’t, careful review is priced as a personal cost.
They don’t add up, they feed each other — which is why oversight rarely fails loudly. It reports success right up until the day it matters.
- Everything gets escalated to review
- Attention per item drops
- Approval becomes the default
- Almost nothing is ever flagged
- The gate looks like it is working, so more is sent through it
Fix the design, not the person. The person is responding rationally to what you built.
The patterns that hold up
Four patterns do most of the real work. Each buys something specific and charges for it.
Gate at chain boundaries, not just the end
The instinct is to review the final output. In a multi-step agent that’s the worst place to look — errors compound, and by the end you’re reviewing a fluent summary of a mistake made four steps ago.
The usual design — one gate, at the end
The reviewer sees a plausible email. They cannot see that the search returned nothing and the summariser invented a finding to cover it — laundered two steps ago. No amount of care rescues this: the evidence needed to catch the error was destroyed before the gate.
The fix — a small check where the failure is still obvious
The empty result is caught while it is still, unmistakably, an empty result. No judgement required — it’s a check your code makes for free.
Failures are easiest to catch where they’re still recognisable as failures. A wrong answer gets more convincing with every step that processes it.
Route by risk, not by confidence
The obvious move is to send the model’s least confident answers to a human. It works less well than it sounds, because a confidence score doesn’t track whether the answer is right — the same mechanism behind a confident wrong answer generally.
Switch the rule below and watch the shaded band.
Interactive: eighty outputs plotted by the model's confidence against the cost of being wrong, routed first on confidence and then on consequence. Needs JavaScript.
Cutting on confidence leaves high-consequence mistakes auto-approved, because the model was sure about them. Cutting on consequence clears that band and sends far less to a human. Confidence can be one input; it shouldn’t be the gate.
- the action cannot be undone
- the model's confidence is below 0.8
- it affects more than one person
- the output is longer than usual
- it touches a regulated field
- no source was retrieved
- the model used more tokens than average
- the output contradicts its source
Make review a workflow state, not a signal
An approve/reject that only produces a quality metric is wasted. Make the decision drive the pipeline: pending blocks, approved proceeds, rejected routes somewhere a person owns.
That gives you the audit trail you’ll need anyway — and it makes “rejected” a real outcome with a real path, which is what separates a check from a step.
Feed corrections back
Every human correction is labelled data describing exactly where your system fails. Capture the correction, not just the verdict. Over time that shrinks the escalation volume, which is the only sustainable way to keep review attention high.
Design the escalation, not just the gate
This is where rubber-stamping is prevented, and it gets the least attention.
Most escalations carry the output and nothing else, so the only available heuristics are “does it read fine” and “am I behind”. Add the layers one at a time.
Interactive: a £4,200 refund approval request, built up one layer at a time — why it escalated, what the agent did, where it's unsure, and what it costs if wrong. Needs JavaScript.
Same decision, completely different cognitive task. The built-up version tells the reviewer where to look — day 15 against a 14-day window — and what it costs to be wrong. It also surfaced its own weakest link instead of hiding it inside fluent prose.
What the law already requires
If you’re in the EU or selling into it, this stops being a design preference. Read properly, Article 14 is a free specification.
Article 14(4) says the people overseeing a high-risk system must be able to:
- understand its capacities and limitations, and monitor it for anomalies
- correctly interpret its output
- decide not to use it, or to “disregard, override or reverse” the output
- interrupt it via a stop button or similar, bringing it “to a halt in a safe state”
And then it names the exact failure mode this post is about. Overseers must be enabled:
to remain aware of the possible tendency of automatically relying or over-relying on the output produced by a high-risk AI system (automation bias)
The law already knows a human in the chair isn’t enough. Your design has to counter the reflex to defer.
The deadlines moved in 2026, so check which one applies to you.
The compliance timeline, and what it doesn't cover
The dates shifted. A 2026 amendment (Regulation (EU) 2026/1744, in force 27 July 2026) pushed the high-risk deadlines back. Standalone high-risk systems now have until 2 December 2027. AI built into products that are already regulated — medical devices, machinery — has until 2 August 2028.
Article 50 was not deferred. The transparency duties, including disclosure and marking of AI-generated content, run on their original schedule. A deferral of high-risk obligations is not a general reprieve.
“High-risk” is a defined category, not a vibe. Most products aren’t in it. Annex III lists the areas — employment, education, essential services, law enforcement, and others. Check whether you’re actually in scope before you build a compliance programme, and check whether you’re a provider or a deployer, because the duties differ.
Verify before you rely on this. These dates changed once already, this post is not legal advice, and it’s written from practice rather than law. If a date is load-bearing for a real decision, read the Official Journal or ask counsel — don’t trust a blog post, including this one.
Outside the EU, the NIST AI Risk Management Framework names human oversight as a process to be defined, assessed and documented — not merely present. It’s voluntary, but it’s the reference most US enterprise procurement now points at.
Article 14 reads like compliance language, but every clause in it names something you build.
- Correctly interpret the output
- Disregard, override or reverse the output
- Interrupt it and bring it to a safe halt
- Stay aware of the tendency to over-rely on it
Either way, oversight is becoming something you evidence, not assert.
How to tell if your oversight is real
You cannot evidence it by pointing at the approval gate. Here’s what does.
Seed known-wrong outputs. Inject items you know are wrong into the review queue and measure how many get caught. It’s the one number that tells you whether oversight is real, and a rate near zero means your reviewers are processing, not reviewing. You’d rather learn that from a seed than from a customer.
How to run a seeded-error test
Build ten to twenty seeds, not two. A handful tells you nothing — you need enough to see a rate. Mix them: some obviously wrong, some subtly wrong. The subtle ones are the real measurement.
Derive them from failures you’ve already had. Your incident log is the best source of realistic wrong answers. Invented errors tend to be too easy, because you unconsciously make them detectable.
Inject them into the normal queue, indistinguishable from real work. A seeded item that arrives flagged, batched or at a suspicious hour tests nothing. If reviewers can spot the test, you’re measuring their test-detection skill.
Tell people the programme exists, but not which items are seeds. Hiding it completely reads as a trap and costs you the team’s trust, which is what the whole thing depends on. Say plainly that it measures the system, not the person. Make it a performance metric and reviewers will start hunting for seeds instead of doing the job.
Record the catch rate, and the time-to-catch. A seed caught after the action fired is not a catch.
Re-run after every change to the model, the prompt, the routing thresholds or the escalation payload. The rate is only meaningful as a trend. A redesigned payload that moves catch rate from 20% to 70% is the clearest evidence you’ll ever get that a design change worked.
One warning: if the catch rate is low, resist the urge to tell reviewers to try harder. That’s the response that produces a worse system and a demoralised team. Low catch rate is a volume, routing or payload problem almost every time.
Watch how often reviewers say no. Approve everything and the gate does nothing. Reject most of it and you’re sending work that should never have reached them.
Measure dwell time per item. A median review of four seconds means nobody read the payload. That’s a design verdict, not a performance-management one.
Test the stop path. Can a reviewer halt the system, and has anyone tried recently? An untested emergency control is decoration — and under Article 14, a gap.
- Pull ten to twenty realistic wrong answers from your incident log
- Tell the team the programme exists, without saying which items are seeds
- Inject them into the normal queue, indistinguishable from real work
- Record the catch rate and the time-to-catch
- Re-run it after every change to the model, routing or payload
Same discipline as building an eval: pick a number, change one thing, see if it moves.
Match the pattern to the stakes
Two questions do most of the deciding: can you undo it, and how many people does it touch?
| Action | Pattern | Why |
|---|---|---|
| Reversible, single user | Out of the loop; log and sample | A wrong answer costs one correction. Review would cost more than the errors |
| Reversible, many users | On the loop, with alerting | Needs monitoring, not a gate — catch the pattern, not each instance |
| Irreversible, bounded cost | In the loop, approval gate with a rich payload | The classic case: money, messages, deletions |
| Irreversible, wide blast radius | Two-person rule, plus a tested stop | One reviewer is one automation-bias failure away from a very bad day |
| Regulated or safety-critical | In the loop, plus audit trail and documented process | Article 14 territory — you’ll need to evidence it, not just do it |
The skill is spending oversight where a wrong answer actually costs something. Escalate everything and you’ve built the 49–96% problem yourself.
And a system designed to refuse when it’s out of its depth never reaches a reviewer at all. Every “I don’t know” is an escalation you didn’t have to staff.
- Reversible, one person
- Reversible, many people
- Irreversible, bounded cost
- Irreversible, wide or regulated
- Renaming a file in a user's own workspace
- An agent emailing 8,000 customers at once
- Issuing a £300 goodwill credit
- A recommendation feed shown to every visitor
- Changing a patient's medication record
- Auto-tagging a photo in a private album
Where oversight usually breaks
- Oversight only at the end. By the final step the error is fluent. Gate at boundaries.
- Unbounded review volume. Attention per item is a budget. Escalating everything spends it on nothing.
- Using the model’s confidence score to decide. It doesn’t track whether the answer is right. Route on consequence instead.
- Bare escalations. “Approve? [y/n]” with no reasoning asks for a signature, not a review.
- A reviewer with no authority to refuse. If “no” creates a problem for them, you’ve built a step.
- No tested stop. An emergency control nobody has pulled is a hypothesis.
- Never testing the reviewer. Seed errors. An untested control isn’t a control.
Where to start
Pick your highest-consequence automated action. Write down what happens if it’s wrong and whether it can be undone — that picks your pattern from the table above.
Then three things. Limit what gets escalated, so attention stays real. Make each escalation carry its reasoning, so there’s something to check against. Seed ten known-wrong items and find out whether anyone catches them.
That last one usually settles the argument.
The goal was never a human in the loop. It was a system that fails safely, and a person genuinely able to catch it when it does.
Check yourself.
- What’s the difference between a review step and a review check?
- The 49–96% override finding came with a nuance about justified overrides. What does that reframe the problem as?
- Which reviewer profile is most susceptible to automation bias, and why is that awkward?
- Why is model confidence a poor signal for deciding what to escalate?
Anything that won’t come is your reread map.
Want the simple end of this running? Sprout refuses rather than guessing, and Scout runs behind a step limit and a daily cap, so its loop stops whether or not anyone is watching. Both are in the Greenhouse — and both are oversight done by design rather than by rota. 🌱