Practice Placing Checks in an LLM Application Request Lifecycle
A schema validator turns green, the response parses, and the builder ships it. The answer is still wrong. The check passed because it was never asked to…

Key topics
A schema validator turns green, the response parses, and the builder ships it. The answer is still wrong. The check passed because it was never asked to prove the thing the builder assumed it proved.
That gap — between what a check establishes and what you wish it established — is the whole subject of this exercise. You will map proposed checks onto five lifecycle boundaries, write down the evidence each one actually produces, run a checker against your map, and find the boundary you left uncovered.
Why Checks Feel Interchangeable Until They Aren't
Most builders start with a mental model that treats validation as a generic safety layer. You sprinkle it on top of the pipeline, like a filter at the end of a drain. Under that model, any check that passes is good news, and more checks are better than fewer.
The stronger model is narrower and more useful: a check is a claim about one boundary, and every boundary proves less than it looks like it does.
A JSON-schema-valid response is parseable and structurally conformant. It is not necessarily true, current, or authorized. A tool call with well-formed arguments matches the tool's contract. It does not mean the side effect is desirable or reversible. A retrieval result that scores high on relevance matched the query. It does not mean the corpus contained the right passage in the first place.
Production incidents live in that space between evidence and guarantee. The check ran. The check passed. The system still failed, because nobody asked what the check was actually capable of establishing.
This exercise makes that question concrete. You will produce a map, run a checker against it, and find the boundary you left empty.
The Five Boundaries You Are Mapping To
If you have already traced the full request lifecycle — input, context assembly, model call, tools, verification — you have the stage order. Here we only need to fix the boundary names the exercise uses, because consistent labels are what make the mapping gradeable.
| Boundary | What it covers | What a check here can establish |
|---|---|---|
| Input | What the user or upstream system supplied, before any model sees it | Shape, size, encoding, and basic admissibility of the request |
| Retrieval | What evidence was selected for this request | Relevance, freshness, and authorization of the selected passages |
| Model output | What the model produced, treated as untrusted data | Structure, schema conformance, and format |
| Tool execution | What the system is about to do in the world | Argument validity, permissions, and side-effect scope |
| Final verification | Whether the assembled result is acceptable to return | End-to-end acceptability of the whole response |
Here is the distinction the table cannot carry on its own: lifecycle placement decides when a check can act; the inputs it inspects and the rule it applies decide what it can establish. Moving a check later in the pipeline does not widen its view. A schema validator that only ever reads the model's JSON stays a schema validator whether you run it at the model output boundary or at final verification. It still cannot tell you whether retrieval selected the right passage or whether a tool call is safe to execute.
So when you map a check, you are answering two separate questions. Where does it run? And what does it actually look at? A check placed at final verification that inspects only one field is still a local check. A check placed early that inspects the full request trace can still make a broad claim. Position and scope are related, but they are not the same thing — and confusing them is the most common way builders overestimate their coverage.
Knowledge check
Check your understanding
Answer this question before you continue.
Set Up the Exercise Files and Run the Checker
Dependencies are minimal on purpose: Python 3 standard library only. No packages, no API keys, no network calls. The whole exercise runs offline and deterministically.
You need three supplied files, provided alongside this article in the exercise bundle:
lifecycle.json— the boundary definitions and stage orderscenarios.json— the request scenarios and their proposed checksmap_checks.py— the checker
You write one file: answers.json. It is a JSON array — one object per proposed check in scenarios.json. Each object maps the check to a boundary and carries two free-text fields: what evidence the check establishes, and what it cannot prove.
[
{
"check_id": "c3",
"boundary": "model_output",
"establishes": "response parses and matches the declared schema",
"cannot_prove": "the values are true, current, or authorized"
}
]
The snippet above is a fragment, not a complete answer file. Your answers.json needs an entry for every check_id that appears in scenarios.json, or the checker will report the missing ones as unmapped.
Then run the checker:
python map_checks.py lifecycle.json scenarios.json answers.json
Read the output rather than skimming it. Expect two kinds of feedback: per-check pass/fail on boundary placement, and a list of uncovered gaps — scenarios where no check is assigned to a boundary that matters. A representative run looks like this:
check c1 boundary=input PASS
check c2 boundary=retrieval PASS
check c3 boundary=model_output PASS
check c4 boundary=tool_execution PASS
uncovered gaps:
scenario s2: final_verification has no assigned check
scenario s4: retrieval has no assigned check
The checker is deterministic because the mapping is a classification problem. Placement can be graded objectively. The reasoning fields are yours to defend, and that is where the real learning happens.
Reading the Feedback: Evidence Versus Guarantee
A correct boundary placement can still carry a wrong evidence claim. The checker flags placement; the reasoning fields are where you find out whether you actually understand the check.
Here is the pattern to internalize, boundary by boundary:
| What this check establishes | What it cannot prove |
|---|---|
| Input is well-formed and within limits | The request is legitimate or well-intentioned |
| Retrieved passages match the query | The corpus contained the right passage at all |
| Output is parseable and schema-conformant | The values are true, current, or authorized |
| Tool arguments match the tool contract | The side effect is desirable or reversible |
| The assembled result is acceptable to return | That it will be acceptable tomorrow, or to a different user |
The recurring trap is treating a boundary check as a system guarantee. Green at input, green at retrieval, green at output — so you skip final verification. That is exactly the sequence that ships a confident, well-structured, wrong answer.
Common mistake: Assuming that because every stage passed its own check, the end-to-end result is verified. Each stage check is a local claim. Only a check that actually inspects the assembled result — and is scoped to judge it — can make a claim about the whole.
Knowledge check
Check your understanding
Answer this question before you continue.
Find the Documented Boundary Gap
The scenarios are constructed so at least one boundary has no check assigned. The checker reports it as an uncovered gap, naming the scenario and the boundary — not the fix. Deciding whether the gap is a real risk or an acceptable omission is your job.
Two gap types matter, and they are not the same:
- A missing check — nothing runs at that boundary.
- A misplaced check — something runs, but at the wrong boundary, so it proves the wrong thing.
The second is more dangerous because it looks like coverage. A check sitting at the model output boundary that was meant to validate retrieval will pass every time and catch nothing.
Do not assume the gap at final verification is always the serious one. Sometimes the expensive gap is at retrieval, where a bad selection poisons everything downstream — the model reasons correctly over the wrong evidence, and every later check confirms the reasoning while missing the poison.
Not every gap needs a new check. Some need a documented assumption, a human review step, or a narrower scenario scope. The judgment is yours; the checker only tells you where the hole is.
Knowledge check
Check your understanding
Answer this question before you continue.
Revise the Map: Move a Check, Change a Scenario
Reading the feedback teaches you to interpret. Revising the map teaches you to predict. Do both modifications.
Modification A: Move one check to a different boundary. Before you run the checker, write down what you expect it to report.
Modification B: Alter a scenario so a previously sufficient check becomes insufficient. Observe which gap appears.
Then re-run the same command against the revised answers.json and diff the feedback against the first run.
Watch for these debugging signals:
- A check that passes placement but fails the evidence field — you classified it right and described it wrong.
- A scenario that reports two gaps after your edit — your change removed coverage you did not notice.
- A boundary that now has redundant checks — you added coverage where it already existed and left the real gap open.
This matters beyond the exercise. In a real system, moving a check changes cost, latency, and blast radius. A check at the tool boundary is far more expensive to get wrong than one at the input boundary, because by the time you reach tool execution, the system is about to act in the world.
Knowledge check
Check your understanding
Answer this question before you continue.
Turning the Map Into a Habit
The exercise is a stand-in for a decision you make every time you add a check to a real pipeline. The rule is short:
Before adding a check, name the boundary, name the evidence, and name what it cannot prove. If you cannot fill all three, the check is decoration.
A few heuristics that follow from the mapping:
- Put the cheapest check that can still catch the failure as early as possible — but never let an early check stand in for final verification.
- Watch for check sprawl: many overlapping checks that all prove the same thing while one boundary stays empty.
- Sometimes the right move is not a check at all. Narrow the scope, add human review, or refuse the request outright.
I have watched builders add a fifth validator to a boundary that already had three, while retrieval ran with no check whatsoever. The instinct to add checks is not the same as the judgment to place them.
Where to Go Next
Take one real request from a pipeline you already own. Map it against the same five boundaries. Count how many have no check at all — not a weak check, not a misplaced one, but nothing.
That count is the gap you have been shipping with. The exercise gave you the vocabulary and the checker. Your own pipeline gives you the stakes.
Knowledge check
Final check
Finish the article by checking the ideas you just learned.
References
Want a more structured LLMOps path?
Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.
Large Language Models Starter Pack
A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.
- 227-page Illustrated PDF edition
- 12 guided LLM engineering chapters
- Visual concept diagrams
- Self-assessment quizzes
- Bonus deep-dive sections
- Prompt design, structured output, context windows & RAG pipelines
- Agents, tool calling, prompt injection, evaluation & application lifecycles
Coming soon


