Practice Deciding When an LLM Output Needs Human Review
The output is sitting in front of you. It reads well. It answers the question. And you still cannot say whether it should ship.

Key topics
The output is sitting in front of you. It reads well. It answers the question. And you still cannot say whether it should ship.
That hesitation is the skill gap this exercise targets. You already know the four factors — uncertainty, impact, reversibility, evidence. You have seen the expected-loss argument for placing review at a decision boundary. What you have not done is commit to a single disposition for a single output and defend it in writing. That is what we are going to practice.
By the end, you will have a justified disposition for every case in a supplied set, a checker run that compares your judgment against a written policy, and one deliberate threshold change that shows how the same output flips disposition when the line moves.
Why Deciding Feels Harder Than Explaining
Reciting the four factors is easy. Applying them to one concrete output is not, because the factors do not resolve themselves into a verdict. Uncertainty is a range. Impact depends on who absorbs the consequence. Reversibility depends on timing. Evidence strength depends on what you can actually point to.
When those four signals conflict, most people fall back on a feeling: looks fine to me. That is not a disposition. It is an unstated policy — one you have never written down, never tested, and cannot defend when someone asks why this output shipped and the last one did not.
A disposition is a commitment. It says: given what I can see, this output goes here, and here is why. The exercise forces that commitment before you get any feedback, because feedback you receive before committing is not practice — it is rationalization with extra steps.
The four dispositions you will use:
| Disposition | What it claims |
|---|---|
| Proceed | Residual error is acceptable for this output in this context |
| Review | A human inspection before release is cheaper than the expected error |
| Reject | The output is not usable and no review path rescues it |
| Escalate | The decision exceeds your authority or your evidence |
Success here is not a high match rate against the policy. Success is a justified disposition for every case — one you can explain without gesturing at a feeling.
The Four Dispositions and What Each One Costs
Before you touch the files, get the vocabulary precise. Beginners routinely collapse reject and escalate into one bucket, and that collapse hides a real distinction.
Proceed means the output ships as-is. The claim is that the residual error rate is acceptable here — not everywhere, not forever, but for this output in this context.
Review means a human inspects before release. The claim is that inspection costs less than the expected loss from shipping unchecked. Review is not free; it consumes reviewer time and attention, and it can be wrong.
Reject is a quality verdict. The output is not usable, and no amount of human inspection rescues it. Regenerate, re-retrieve, or drop the task. Rejection is about the artifact.
Escalate is a jurisdiction verdict. The output might be fine, but the decision to ship it is not yours to make — it exceeds your authority, your evidence, or your role. Escalation is about who decides, not what the output says.
That distinction matters because the two failure signals point at different fixes. A reject tells you the generation or retrieval step needs work. An escalate tells you the decision rights need work. Confusing them means you keep regenerating outputs when the real problem is that nobody has decided who owns the call.
These four map onto the expected-loss reasoning you have already seen: proceed when expected loss is below the review cost, review when it is above, reject when no disposition recovers the output, escalate when the loss calculation itself is outside your scope. We are not re-deriving that model here. We are applying it.
Knowledge check
Check your understanding
Answer this question before you continue.
Set Up the Exercise
You need Python 3 and nothing else. No installs, no API keys, no network access. The standard library is sufficient.
Three files are supplied:
cases.json— the LLM outputs you will judge, each with fields for uncertainty, impact, reversibility, and evidencepolicy.json— the thresholds and rules that define the baseline disposition for each combinationreview_decisions.py— the checker that compares your answer sheet against the policy
Your job is to produce a fourth file: answers.json, one entry per case, with a disposition and four short reasons.
The run command:
python review_decisions.py cases.json policy.json answers.json
Observable success looks like per-case feedback lines showing whether your disposition matched the policy-derived disposition, plus the rule that fired. Mismatches are not failures. They are the interesting output.
Read the Policy Before You Read the Cases
Here is the discipline that separates practice from guessing: read policy.json first.
The policy encodes thresholds over the four factors. A threshold turns a fuzzy judgment into a comparable category. "High uncertainty" becomes a boundary you can check against. "Weak evidence" becomes a condition the checker can evaluate.
If you read the cases first, you will anchor on gut feel and then fit your reasons to your preferred disposition. You will not notice yourself doing it. The policy is the baseline; your instinct is the hypothesis under test. Read the rule set before you meet the outputs.
A rough decision flow:
case fields
→ compare uncertainty against threshold
→ compare impact against threshold
→ check reversibility
→ check evidence strength
→ one of: proceed | review | reject | escalate
The policy is authored, not discovered. Someone chose those thresholds. That choice is the real design decision, and you will test it shortly.
Knowledge check
Check your understanding
Answer this question before you continue.
Fill In Your Answer Sheet
Now the core drill. Here is the reasoning shape for one case.
Suppose a case shows moderate uncertainty, high impact, irreversible consequences, and weak evidence. Read the fields. Compare each against the policy threshold. Commit to a disposition. Then write one line per factor explaining why that factor pushed you where it did.
A good reason names the consequence and who absorbs it. "High impact" alone is not a reason. "High impact because a wrong answer here reaches the customer directly and cannot be recalled" is a reason.
Evidence strength is separate from model confidence. A confident output with no supporting evidence is still weak evidence. The model's tone is not a signal about correctness; the presence of verifiable support is.
Your answers.json needs one entry per case. The shape is a list of objects, each keyed to a case and carrying your disposition plus the four reasons:
[
{
"case_id": "case_001",
"disposition": "review",
"uncertainty_reason": "Model gave two plausible readings of the same clause.",
"impact_reason": "A wrong answer reaches the customer directly.",
"reversibility_reason": "Once sent, the message cannot be recalled.",
"evidence_reason": "No source document supports the specific claim."
}
]
Map one supplied case into that shape before you write the rest. Open cases.json, find the first case, and copy its case_id exactly. Then read its four factor fields and write one reason per factor. The checker matches on case_id, so a typo there produces a missing-case error rather than a disposition mismatch — a different debugging signal worth recognizing.
Common mistake: Writing your reasons after you see the checker's feedback. That converts practice into rationalization. Commit first, run second.
Work through every case before you run anything. Write the disposition. Write the four reasons. Then move on.
Knowledge check
Check your understanding
Answer this question before you continue.
Run the Checker and Read the Feedback
Now run it.
python review_decisions.py cases.json policy.json answers.json
The checker compares your disposition against the policy-derived disposition, per case. Expected output shape: per-case lines with your disposition, the policy disposition, and the rule that fired.
Mismatches come from two distinct causes, and you need to tell them apart:
- You misread the case fields. You treated moderate uncertainty as low, or missed that the consequence was reversible. This is a reading error, and it is fixable.
- You disagree with the policy. You read the fields correctly and still think the threshold is wrong for this context. This is legitimate — but it must be stated as a policy critique, not a vibe.
"I felt like it should be reviewed" is not a critique. "The policy treats this impact level as review-worthy, but the consequence here is recoverable within the same session, so proceed is defensible" is a critique. Record it. Disagreement with a written policy is useful information about where the policy needs work.
Change One Threshold and Rerun
This is the experiment that makes the lesson stick. Open policy.json and move one threshold — raise the impact boundary that triggers review, or lower the evidence bar that forces escalation. Rerun the same command with the same answers.json.
Watch which cases flip disposition and which stay put. The cases that flip are the ones sitting near the line you moved. The cases that hold are the ones whose factors were decisive regardless of that threshold.
This is the point: the same output can be proceed or review depending on where the line sits. The output did not change. The line did. That means the line is the real design decision, and the output is downstream of it.
Explain the change in one sentence: which threshold moved, which cases crossed it, and why that is the correct consequence. If you cannot explain why the flip is correct, you have found a threshold you do not actually understand yet.
If you also want a sanity check on the checker itself, flip one of your own answers and rerun. That should produce a mismatch on exactly that case and nothing else. It tests whether the checker is reading your answer sheet correctly — it does not change the policy-derived disposition, so it is a different experiment from moving the threshold.
When the Script Breaks: Malformed JSON and Missing Fields
The checker will refuse to run on a broken answer sheet, and that refusal is a debugging signal, not an obstacle.
Malformed JSON usually means a trailing comma, an unquoted key, or a stray comment. Read the parser error and find the line number. JSON does not forgive.
Missing fields mean the checker cannot evaluate a factor. Without a reason for uncertainty, the disposition is unverifiable. A missing reason field is a real failure, not a formatting nit — an unjustified disposition is exactly the thing this exercise exists to prevent.
Warning: Do not patch the checker to accept incomplete answers. Fix the answer sheet. The constraint is the point.
The debugging loop is short: read the error, locate the field, fix the input, rerun.
Knowledge check
Check your understanding
Answer this question before you continue.
What This Exercise Does Not Prove
A passing run against a toy policy is not evidence that a real workflow is safe. Be honest about the boundary.
The policy is authored, not discovered. A wrong threshold produces confidently wrong dispositions, and the checker will happily confirm them.
Synthetic cases are cleaner than production outputs. Real outputs arrive with missing context, ambiguous provenance, and fields nobody filled in.
The checker validates consistency with a policy. It does not validate correctness of the underlying output. A proceed that matches the policy is still a proceed on an output that might be wrong.
Reviewer error and reviewer cost sit outside this model. A review disposition is not a guarantee — it is a bet that inspection catches what shipping would miss.
Where to Take This Next
The rule to carry: never ship an LLM output you cannot justify against a written policy. If you cannot name the threshold that let it through, you did not decide — you defaulted.
Treat every mismatch between your instinct and the policy as a signal to inspect one of the two. Either your reading was wrong, or the policy is. Both are worth knowing.
The next move is your own workflow. Sample ten real outputs. Write the answer sheet. See which dispositions you cannot defend. The ones you cannot defend are the ones where your policy is missing, and that is where the real work starts.
Knowledge check
Final check
Finish the article by checking the ideas you just learned.
References
Want a more structured LLMOps path?
Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.
Large Language Models Starter Pack
A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.
- 227-page Illustrated PDF edition
- 12 guided LLM engineering chapters
- Visual concept diagrams
- Self-assessment quizzes
- Bonus deep-dive sections
- Prompt design, structured output, context windows & RAG pipelines
- Agents, tool calling, prompt injection, evaluation & application lifecycles
Coming soon


