When Is Human Review Worth It? Derive an Expected-Loss Threshold
A review queue is a purchase order. You are buying a reduction in expected loss, and the only real question is whether the reduction is worth the price.

Key topics
A review queue is a purchase order. You are buying a reduction in expected loss, and the only real question is whether the reduction is worth the price.
Most teams ship the queue first and argue about the cutoff second. Someone says "escalate below 0.9 confidence," nobody can say where 0.9 came from, and the queue either drowns the reviewers or lets the expensive errors through. The number feels principled because it is numeric. It is not a decision rule.
This article builds the decision rule. We will derive a small expected-loss threshold you can write on a whiteboard, argue about with a teammate, and tune with real numbers. The earlier piece on where review belongs established the qualitative boundary — uncertainty, impact, reversibility, verification. This one prices it.
Why a Confidence Cutoff Is Not a Decision Rule
The default pattern is familiar: pick a confidence number, route everything below it to a human. It breaks for two reasons.
First, confidence is a property of the model's output, not of the decision. A 0.85 on a reversible draft and a 0.85 on an irreversible payment are not the same event. The cutoff treats them as identical because it only sees one number.
Second, confidence is frequently miscalibrated. A model that says 0.9 may be right far less often than 90% of the time, which means your cutoff is anchored to a number that does not mean what you assume. The threshold you set and the error rate you get can drift apart without anyone noticing.
The replacement frame is simpler: human review is a purchase. You pay a known cost to buy a reduction in expected loss. Buy it when the reduction is larger than the price. That reframes the whole argument from "what confidence is safe?" to "what is this review actually buying?"
Four quantities drive the answer: the probability the output is wrong, the impact if that error reaches the world, the cost of one review, and the residual error that survives review. Everything below moves those four around.
Knowledge check
Check your understanding
Answer this question before you continue.
Notation and Assumptions Before Any Formula
Define each symbol in plain language before using it, because the derivation is only as honest as its inputs.
- p — the probability the automated output is wrong on this case. This must be a calibrated estimate, not a raw model score. If your confidence signal is uncalibrated, p is a guess wearing a decimal point.
- C — the impact if that error reaches the world unchecked, expressed in one consistent unit: dollars, minutes, tickets, or a normalized severity score.
- R — the fully loaded cost of one review. Reviewer time, latency added to the workflow, and the friction of a queue all belong here.
- r — the residual error probability after review. The reviewer misses it, or the reviewer is wrong in the other direction.
State the assumptions explicitly, because each one is a simplification, not a truth:
- One decision per case.
- A scalar impact — a single number captures the damage.
- Review cost is independent of the case.
- p is calibrated.
Common mistake: Mixing units. If C is in dollars and R is in hours, the model is nonsense. Pick one unit and stay in it for the whole calculation.
Deriving the Review Threshold Step by Step
Build the comparison from two expected-loss expressions rather than presenting a formula wall.
Step 1 — Expected loss if you do not review. The error happens with probability p, and costs C when it does:
L_no_review = p × C
Step 2 — Expected loss if you do review. You pay R whether or not the reviewer finds anything, plus the residual loss r × C:
L_review = R + (r × C)
Step 3 — Set them against each other. Review is worth it when the unchecked loss exceeds the reviewed loss:
p × C > R + (r × C)
Step 4 — Isolate impact. Move the residual term to the left and factor out C:
C × (p − r) > R
C > R / (p − r)
Name the right-hand side the threshold. Cases with impact above it get reviewed; cases below it do not.
Step 5 — Read the result out loud. The denominator, p − r, is the error reduction the review actually buys. If review barely reduces error, the denominator collapses and the threshold explodes. A review that catches almost nothing has to be justified by enormous impact to pay for itself.
Degenerate cases. If p ≤ r, review is never worth it on expected-loss grounds alone, and the threshold is undefined or infinite. You are paying R to make things no better, or worse.
Picture a single line representing impact. The threshold sits somewhere on it. Review lives to the right; automate lives to the left.
Knowledge check
Check your understanding
Answer this question before you continue.
A Worked Example With Real Numbers
Take a support-triage case: an automated routing decision where a misroute costs a delayed customer and a wasted specialist hour.
| Symbol | Meaning | Value |
|---|---|---|
| p | probability of a wrong route | 0.08 |
| r | residual error after review | 0.02 |
| R | reviewer time per case | $12 |
| C | downstream impact of a misroute | $400 |
Compute the threshold:
C > R / (p − r)
C > 12 / (0.08 − 0.02)
C > 12 / 0.06
C > 200
Cases above $200 of impact get reviewed. Below $200, they do not.
Interpret the number rather than just reporting it. The review buys a 6-point reduction in error probability, from 8% to 2%. At $12 per review, each point of error reduction costs $2. The threshold is where that $2-per-point price meets the impact at stake.
Now watch the sensitivity. The threshold is a ratio, so it moves fast when the inputs move.
| Change | New threshold |
|---|---|
| Halve R to $6 | $100 |
| Double the error reduction to 12 points | $100 |
| Drop error reduction to 1 point | $1,200 |
Halve the review cost and the threshold halves. Double what the review buys and the threshold halves again. Shrink the error reduction to a single point and the threshold jumps sixfold. The denominator is doing most of the work.
How Impact, Uncertainty, and Reversibility Move the Threshold
The four inputs are not equally trustworthy, and one of them is missing from the base model entirely.
Impact is the hardest number to estimate and it dominates the result. A wrong C swamps a careful p. If you spend an afternoon calibrating error rates and guess the impact, you have optimized the cheap term.
Uncertainty in p matters most near the threshold. If p is 0.08 plus or minus 0.04, the threshold is a range, not a point. Treat it that way. A cutoff that flips between $100 and $600 depending on which sample you drew is not a policy yet.
Reversibility is not in the base model. Introduce a severity multiplier s between 0 and 1, where 1 means irreversible and small values mean the error is usually caught downstream. The extended condition becomes:
C > R / ((p − r) × s)
Reversible errors tolerate much higher impact before review is justified, because much of their damage gets caught and corrected later.
Run the same numbers with s = 0.1. The threshold moves from $200 to $2,000. Same model, same reviewer, ten times the tolerance — because the error is mostly self-correcting.
Warning: Silent errors that nothing downstream catches behave like
s = 1even when they look small. That is why they deserve a lower threshold than their dollar value suggests. The absence of a safety net is itself a cost.
Knowledge check
Check your understanding
Answer this question before you continue.
When the Reviewer Is Also Wrong
Residual error is a first-class term, not a footnote. r depends on reviewer expertise, time pressure, queue depth, and how legible the model's output is.
Reviewer fatigue and inconsistent rubrics make r drift upward over a long shift. As r rises, the denominator p − r shrinks, the threshold climbs, and more errors slip through — quietly, because nobody re-derived the cutoff when the reviewers got tired.
If r approaches p, review becomes theater. You pay R and buy almost nothing. The queue looks like a control; it is a receipt.
This is the argument for measuring review quality directly rather than assuming it. Track disagreement between reviewers and the original labels instead of treating historical labels as ground truth. Automated checks and periodic human review serve different jobs, and review is most valuable where automated checks are weakest.
Tip: If you cannot estimate r at all, you cannot claim the review is worth its cost. Measure it on a sample before you touch the cutoff.
Knowledge check
Check your understanding
Answer this question before you continue.
Where the Model Stops Being Useful
A deliberately small model has honest boundaries. Know them before you over-apply it.
- It assumes independent cases. It breaks when errors correlate, when one bad output poisons many downstream decisions, or when the queue itself becomes the bottleneck.
- It assumes a scalar impact. Real errors have legal, reputational, and trust dimensions that do not collapse cleanly into one number.
- Some decisions are not priced at all. Regulatory sign-off, contractual approval, and policy-mandated review happen because a rule requires them, not because the arithmetic favors them. Do not pretend the model overrides those.
- It assumes review cost is constant. In practice R rises with queue depth, so a threshold that sends too many cases to review can make each review more expensive and less accurate at the same time.
- It assumes you can estimate p. If your confidence signal is uncalibrated, the whole derivation rests on a number you have not validated.
Use the model to structure the argument and expose which assumption is doing the work — not to produce a single authoritative cutoff.
Turning the Threshold Into a Working Policy
Here is the procedure I would run on a real system this week.
- Write down your four numbers before you argue about the cutoff. Most review debates are actually disagreements about C or r, not about the threshold.
- Pick one unit and one time window, and state both in the policy so the number stays interpretable when someone revisits it.
- Set the threshold on a tuning set, then check it on held-out cases. A threshold chosen on the same data you evaluate on will look better than it is.
- Re-derive when the model version, the label definitions, or the routing policy changes, because p and r both move.
- Log the review rate, the missed high-impact cases, and the reviewer disagreement rate. Those three numbers tell you whether the threshold is doing its job.
The decision rule in one line: review when the expected loss you prevent exceeds the cost of looking.
That rule is only as good as the numbers you feed it, and the two you probably have not measured are p and r. Before you touch the cutoff, instrument your system to estimate both on a sample — then re-derive. The natural next skill is measuring review quality and calibration directly, because a threshold built on an uncalibrated confidence signal is just a guess with better formatting.
Knowledge check
Final check
Finish the article by checking the ideas you just learned.
References
Want a more structured LLMOps path?
Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.
Large Language Models Starter Pack
A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.
- 227-page Illustrated PDF edition
- 12 guided LLM engineering chapters
- Visual concept diagrams
- Self-assessment quizzes
- Bonus deep-dive sections
- Prompt design, structured output, context windows & RAG pipelines
- Agents, tool calling, prompt injection, evaluation & application lifecycles
Coming soon


