← all posts

The human who approves everything

There’s a reviewer somewhere approving four hundred AI outputs a shift.

They approve almost all of them. They have to — the queue doesn’t stop, and nobody thanks them for being slow.

On the architecture diagram, that person is a control. In reality they’re a formality. The system has human-in-the-loop AI in exactly the way a fire door propped open has a fire door.

A step versus a check Two pipelines compared. In the first, labelled a step, three items flow into a review box and all three come out approved; nothing can be stopped, so the review is a formality. In the second, labelled a check, three items flow into the same review box but one is diverted out and halted while the other two continue; because rejection is a real and supported outcome, the review is genuinely a control. a step the chain pauses here review all of them pass a check the chain can stop here review two pass one is stopped

That’s the whole distinction. It isn’t a people problem — it’s a design problem, and you can measure whether yours works.

In one breath: Adding a reviewer creates a step, not a check. A check is a point where the chain can actually be stopped.

Three forces collapse it: volume, automation bias, and incentives that reward throughput. A 2006 review of clinical alerting found override rates of 49–96% — high-volume, low-precision review destroys itself.

So bound the volume, route by consequence rather than confidence, gate at chain boundaries, and make each escalation carry what the agent tried and why it’s unsure.

Then seed known-wrong outputs and measure the catch rate. Oversight you haven’t tested is a claim, not a control.

Oversight is a design property, not a staffing decision

“We’ll have a human review it” usually ends a risk conversation without changing a system.

A step is a point where the chain pauses and a person does something. A check is a step that can stop the chain — where “no” is a real, supported, consequence-free outcome for the person saying it.

Three questions separate them:

  • Can the reviewer actually say no, and does the system handle that path?
  • Do they have enough context to know when to say no?
  • Is their workload low enough that saying no is possible?

Five setups below. Make the call before you read the reasoning — a guess you’ve committed to is one you remember being wrong about.

Interactive: five review setups to call as a step or a check, with the reasoning for each. Needs JavaScript.

Everything below is how to build the second kind.

The three positions, and what each commits you to

“Human in the loop” covers three genuinely different architectures. The difference is where the person sits relative to the action.

Where the human sits in each of the three oversight positionsIn the in-the-loop position, the AI output goes to a human who decides, and only then does the action happen. In the on-the-loop position, the action happens immediately and a human monitors afterwards, retaining the ability to veto or intervene. In the out-of-the-loop position, the action happens with no human involved at runtime, governed only by boundaries that people set earlier at design time.

AI produces an output

IN the loop

ON the loop

OUT of the loop

Human decides

Action happens

Action happens

Human monitors, can veto

Action happens

Bounds set at design time

PositionWho decidesExecution modelThe failure it invites
In the loopThe human. AI recommends, nothing proceeds without inputSynchronous — the workflow pauses and waitsRubber-stamping under volume; unbounded latency
On the loopThe AI. Humans monitor and hold vetoAsynchronous — dashboards and override surfacesNobody watching the dashboard; discovering failures late
Out of the loopThe AI, inside boundaries set at design timeFully autonomousBoundaries that were wrong, discovered in production

Each position commits you to a different system, not a different policy.

  • In the loop — the workflow stops dead and resumes where it left off. Someone goes to lunch and your system has to survive it.
  • On the loop — a dashboard nobody opens fails the same way a reviewer who approves everything does.
  • Out of the loop — right when a mistake is small and easy to undo. Claiming oversight you don’t have hides the problem somewhere harder to find.
  • IN the loop
  • ON the loop
  • OUT of the loop
  • Tagging support tickets by topic
  • Sending a £4,200 refund
  • A pricing model updating 40,000 listings overnight
  • Deleting a customer's account
  • Autocompleting a search box
  • An agent posting status updates to a team channel

Why oversight collapses: the evidence

“People rubber-stamp AI” gets asserted constantly and evidenced almost never.

The override numbers

Clinical decision support has run this experiment at scale for twenty years: drug-safety alerts fire, a clinician responds, every response is logged.

In 2006 van der Sijs et al. pulled together seventeen studies of it. Clinicians overrode safety alerts in 49% to 96% of cases.

The part usually left out: the same review found overriding is often justified. The alerts fired too readily, on interactions that didn’t matter, to people who knew better.

So the lesson isn’t that clinicians stopped caring. It’s that high-volume, low-precision review destroys itself.

Interactive: two dials — items sent to review per shift, and the seconds a careful review needs — driving the share of flawed work that gets caught. Needs JavaScript.

Drag the volume up and the gate is still there. It just stops doing anything — the reviewer adapts correctly to a badly designed system, and every item becomes noise, including the one that mattered.

If your AI escalates everything, that’s the system you’re building.

Who actually rubber-stamps

Pause & recall. You’re staffing a review queue. Who is most likely to wave through a wrong AI output — someone who has never used AI, someone who uses it daily but isn’t technical, or an AI expert? Commit to an answer before you read on.

Reveal the answer

The one in the middle. Automation bias peaks at low-to-moderate AI familiarity — enough exposure to trust the output, not enough to know how it fails. Complete novices lean towards distrusting it. Experts trust it about as much as it deserves. The awkward part is that the middle profile is the obvious one to staff a review queue with.

The automation bias curve A line chart plotting the tendency to defer to AI output against a person's familiarity with AI. The line starts slightly below the calibrated-trust baseline for people who have never used AI, indicating mild algorithm aversion. It rises steeply to a pronounced peak for people who use AI daily but are not technical, well above the baseline. It then falls away and levels off just above the baseline for people with deep expertise, indicating roughly calibrated trust. calibrated trust peak bias defers to AI output → familiarity with AI → never used AI uses it daily, not technical deep expertise

Horowitz & Kahn tested 9,000 adults across nine countries. They published the method before running it, so the story couldn’t be fitted to the result.

The line rose, peaked, and came back down. The reviewer most likely to defer to a wrong AI output is the moderately AI-familiar one — comfortable with the tools, not deep in how they fail. That profile is the obvious one to put on a review queue.

The three forces, together

How human oversight collapses into rubber-stampingEscalating a high volume of items reduces the attention available for each one. Thin attention, combined with automation bias and with incentives that reward throughput while never measuring catches, makes approval the default response. Because almost nothing is ever flagged, the system appears to be working correctly, which encourages routing even more items through review and reinforces the whole cycle.so route even more

Escalate everything

Attention per item drops

Approval becomes the default

Automation bias

Throughput rewarded, catches not

Nothing flagged, so it looks fine

  • Volume. Attention per item is total attention divided by item count. You control the denominator.
  • Automation bias. Over-weighting machine output against contradicting evidence, documented across aviation, radiology and hiring.
  • Incentives. If throughput is measured and catches aren’t, careful review is priced as a personal cost.

They don’t add up, they feed each other — which is why oversight rarely fails loudly. It reports success right up until the day it matters.

  1. Everything gets escalated to review
  2. Attention per item drops
  3. Approval becomes the default
  4. Almost nothing is ever flagged
  5. The gate looks like it is working, so more is sent through it

Fix the design, not the person. The person is responding rationally to what you built.

The patterns that hold up

Four patterns do most of the real work. Each buys something specific and charges for it.

Gate at chain boundaries, not just the end

The instinct is to review the final output. In a multi-step agent that’s the worst place to look — errors compound, and by the end you’re reviewing a fluent summary of a mistake made four steps ago.

The usual design — one gate, at the end

A pipeline with a single human review gate at the endAn agent searches and returns nothing, the summariser invents a finding to fill the gap, and the draft email reads perfectly. Only then does a human review it before sending. The original failure, an empty search result, has been processed into fluent prose by the time it reaches the reviewer, so there is nothing visibly wrong to catch.

search returns nothing

summarise invents a finding

draft email reads perfectly

human review

send

The reviewer sees a plausible email. They cannot see that the search returned nothing and the summariser invented a finding to cover it — laundered two steps ago. No amount of care rescues this: the evidence needed to catch the error was destroyed before the gate.

The fix — a small check where the failure is still obvious

The same pipeline with an automatic check at the chain boundaryAfter the search step an automatic check asks whether any results were returned. If none were, the chain halts there. Only if results exist does the work continue to summarise and draft, and the human review before sending is then meaningful because the output rests on real retrieved material.noyes

search

any results?

halt here

summarise

draft email

human review

send

The empty result is caught while it is still, unmistakably, an empty result. No judgement required — it’s a check your code makes for free.

Failures are easiest to catch where they’re still recognisable as failures. A wrong answer gets more convincing with every step that processes it.

Route by risk, not by confidence

The obvious move is to send the model’s least confident answers to a human. It works less well than it sounds, because a confidence score doesn’t track whether the answer is right — the same mechanism behind a confident wrong answer generally.

Switch the rule below and watch the shaded band.

Interactive: eighty outputs plotted by the model's confidence against the cost of being wrong, routed first on confidence and then on consequence. Needs JavaScript.

Cutting on confidence leaves high-consequence mistakes auto-approved, because the model was sure about them. Cutting on consequence clears that band and sends far less to a human. Confidence can be one input; it shouldn’t be the gate.

  • the action cannot be undone
  • the model's confidence is below 0.8
  • it affects more than one person
  • the output is longer than usual
  • it touches a regulated field
  • no source was retrieved
  • the model used more tokens than average
  • the output contradicts its source

Make review a workflow state, not a signal

An approve/reject that only produces a quality metric is wasted. Make the decision drive the pipeline: pending blocks, approved proceeds, rejected routes somewhere a person owns.

That gives you the audit trail you’ll need anyway — and it makes “rejected” a real outcome with a real path, which is what separates a check from a step.

Feed corrections back

Every human correction is labelled data describing exactly where your system fails. Capture the correction, not just the verdict. Over time that shrinks the escalation volume, which is the only sustainable way to keep review attention high.

Design the escalation, not just the gate

This is where rubber-stamping is prevented, and it gets the least attention.

Most escalations carry the output and nothing else, so the only available heuristics are “does it read fine” and “am I behind”. Add the layers one at a time.

Interactive: a £4,200 refund approval request, built up one layer at a time — why it escalated, what the agent did, where it's unsure, and what it costs if wrong. Needs JavaScript.

Same decision, completely different cognitive task. The built-up version tells the reviewer where to look — day 15 against a 14-day window — and what it costs to be wrong. It also surfaced its own weakest link instead of hiding it inside fluent prose.

What the law already requires

If you’re in the EU or selling into it, this stops being a design preference. Read properly, Article 14 is a free specification.

Article 14(4) says the people overseeing a high-risk system must be able to:

  • understand its capacities and limitations, and monitor it for anomalies
  • correctly interpret its output
  • decide not to use it, or to “disregard, override or reverse” the output
  • interrupt it via a stop button or similar, bringing it “to a halt in a safe state”

And then it names the exact failure mode this post is about. Overseers must be enabled:

to remain aware of the possible tendency of automatically relying or over-relying on the output produced by a high-risk AI system (automation bias)

The law already knows a human in the chair isn’t enough. Your design has to counter the reflex to defer.

The deadlines moved in 2026, so check which one applies to you.

The compliance timeline, and what it doesn't cover

The dates shifted. A 2026 amendment (Regulation (EU) 2026/1744, in force 27 July 2026) pushed the high-risk deadlines back. Standalone high-risk systems now have until 2 December 2027. AI built into products that are already regulated — medical devices, machinery — has until 2 August 2028.

Article 50 was not deferred. The transparency duties, including disclosure and marking of AI-generated content, run on their original schedule. A deferral of high-risk obligations is not a general reprieve.

“High-risk” is a defined category, not a vibe. Most products aren’t in it. Annex III lists the areas — employment, education, essential services, law enforcement, and others. Check whether you’re actually in scope before you build a compliance programme, and check whether you’re a provider or a deployer, because the duties differ.

Verify before you rely on this. These dates changed once already, this post is not legal advice, and it’s written from practice rather than law. If a date is load-bearing for a real decision, read the Official Journal or ask counsel — don’t trust a blog post, including this one.

Outside the EU, the NIST AI Risk Management Framework names human oversight as a process to be defined, assessed and documented — not merely present. It’s voluntary, but it’s the reference most US enterprise procurement now points at.

Article 14 reads like compliance language, but every clause in it names something you build.

  • Correctly interpret the output
  • Disregard, override or reverse the output
  • Interrupt it and bring it to a safe halt
  • Stay aware of the tendency to over-rely on it

Either way, oversight is becoming something you evidence, not assert.

How to tell if your oversight is real

You cannot evidence it by pointing at the approval gate. Here’s what does.

Seed known-wrong outputs. Inject items you know are wrong into the review queue and measure how many get caught. It’s the one number that tells you whether oversight is real, and a rate near zero means your reviewers are processing, not reviewing. You’d rather learn that from a seed than from a customer.

How to run a seeded-error test

Build ten to twenty seeds, not two. A handful tells you nothing — you need enough to see a rate. Mix them: some obviously wrong, some subtly wrong. The subtle ones are the real measurement.

Derive them from failures you’ve already had. Your incident log is the best source of realistic wrong answers. Invented errors tend to be too easy, because you unconsciously make them detectable.

Inject them into the normal queue, indistinguishable from real work. A seeded item that arrives flagged, batched or at a suspicious hour tests nothing. If reviewers can spot the test, you’re measuring their test-detection skill.

Tell people the programme exists, but not which items are seeds. Hiding it completely reads as a trap and costs you the team’s trust, which is what the whole thing depends on. Say plainly that it measures the system, not the person. Make it a performance metric and reviewers will start hunting for seeds instead of doing the job.

Record the catch rate, and the time-to-catch. A seed caught after the action fired is not a catch.

Re-run after every change to the model, the prompt, the routing thresholds or the escalation payload. The rate is only meaningful as a trend. A redesigned payload that moves catch rate from 20% to 70% is the clearest evidence you’ll ever get that a design change worked.

One warning: if the catch rate is low, resist the urge to tell reviewers to try harder. That’s the response that produces a worse system and a demoralised team. Low catch rate is a volume, routing or payload problem almost every time.

Watch how often reviewers say no. Approve everything and the gate does nothing. Reject most of it and you’re sending work that should never have reached them.

Measure dwell time per item. A median review of four seconds means nobody read the payload. That’s a design verdict, not a performance-management one.

Test the stop path. Can a reviewer halt the system, and has anyone tried recently? An untested emergency control is decoration — and under Article 14, a gap.

  1. Pull ten to twenty realistic wrong answers from your incident log
  2. Tell the team the programme exists, without saying which items are seeds
  3. Inject them into the normal queue, indistinguishable from real work
  4. Record the catch rate and the time-to-catch
  5. Re-run it after every change to the model, routing or payload

Same discipline as building an eval: pick a number, change one thing, see if it moves.

Match the pattern to the stakes

Two questions do most of the deciding: can you undo it, and how many people does it touch?

Choosing a human oversight pattern from reversibility and blast radiusStart by asking whether the action can be undone. If it can, ask how many people it affects: one person means no gate is needed and you should log and sample instead, while many people means monitoring on the loop with alerting. If the action cannot be undone, ask about the blast radius: a bounded cost means an approval gate in the loop with a rich escalation payload, while a wide or regulated impact means a two-person rule with a tested stop control and an audit trail.yesonemanynobounded costwide or regulated

Can the action be undone?

How many people affected?

OUT of the loop — log and sample

ON the loop — monitor and alert

Blast radius?

IN the loop — approval gate, rich payload

Two-person rule, tested stop, audit trail

ActionPatternWhy
Reversible, single userOut of the loop; log and sampleA wrong answer costs one correction. Review would cost more than the errors
Reversible, many usersOn the loop, with alertingNeeds monitoring, not a gate — catch the pattern, not each instance
Irreversible, bounded costIn the loop, approval gate with a rich payloadThe classic case: money, messages, deletions
Irreversible, wide blast radiusTwo-person rule, plus a tested stopOne reviewer is one automation-bias failure away from a very bad day
Regulated or safety-criticalIn the loop, plus audit trail and documented processArticle 14 territory — you’ll need to evidence it, not just do it

The skill is spending oversight where a wrong answer actually costs something. Escalate everything and you’ve built the 49–96% problem yourself.

And a system designed to refuse when it’s out of its depth never reaches a reviewer at all. Every “I don’t know” is an escalation you didn’t have to staff.

  • Reversible, one person
  • Reversible, many people
  • Irreversible, bounded cost
  • Irreversible, wide or regulated
  • Renaming a file in a user's own workspace
  • An agent emailing 8,000 customers at once
  • Issuing a £300 goodwill credit
  • A recommendation feed shown to every visitor
  • Changing a patient's medication record
  • Auto-tagging a photo in a private album

Where oversight usually breaks

  • Oversight only at the end. By the final step the error is fluent. Gate at boundaries.
  • Unbounded review volume. Attention per item is a budget. Escalating everything spends it on nothing.
  • Using the model’s confidence score to decide. It doesn’t track whether the answer is right. Route on consequence instead.
  • Bare escalations. “Approve? [y/n]” with no reasoning asks for a signature, not a review.
  • A reviewer with no authority to refuse. If “no” creates a problem for them, you’ve built a step.
  • No tested stop. An emergency control nobody has pulled is a hypothesis.
  • Never testing the reviewer. Seed errors. An untested control isn’t a control.

Where to start

Pick your highest-consequence automated action. Write down what happens if it’s wrong and whether it can be undone — that picks your pattern from the table above.

Then three things. Limit what gets escalated, so attention stays real. Make each escalation carry its reasoning, so there’s something to check against. Seed ten known-wrong items and find out whether anyone catches them.

That last one usually settles the argument.

The goal was never a human in the loop. It was a system that fails safely, and a person genuinely able to catch it when it does.

Check yourself.

  1. What’s the difference between a review step and a review check?
  2. The 49–96% override finding came with a nuance about justified overrides. What does that reframe the problem as?
  3. Which reviewer profile is most susceptible to automation bias, and why is that awkward?
  4. Why is model confidence a poor signal for deciding what to escalate?

Anything that won’t come is your reread map.

Want the simple end of this running? Sprout refuses rather than guessing, and Scout runs behind a step limit and a daily cap, so its loop stops whether or not anyone is watching. Both are in the Greenhouse — and both are oversight done by design rather than by rota. 🌱

Say hello

Let's grow something together.

Consulting, teaching, speaking, or a product idea that needs an AI brain — my inbox is open.

us — Usama Shahid © 2026 Usama Shahid — reachusama.com 🌱