Skip to content
intermediate

Practice Placing Checks in an LLM Application Request Lifecycle

A schema validator turns green, the response parses, and the builder ships it. The answer is still wrong. The check passed because it was never asked to…

Published 2026-10-03Updated 2026-10-049 min read
Close-up of black and yellow sand patterns on a beach showcasing natural textures.
Close-up of black and yellow sand patterns on a beach showcasing natural textures. Photo by Thilina Alagiyawanna on Pexels.

A schema validator turns green, the response parses, and the builder ships it. The answer is still wrong. The check passed because it was never asked to prove the thing the builder assumed it proved.

That gap — between what a check establishes and what you wish it established — is the whole subject of this exercise. You will map proposed checks onto five lifecycle boundaries, write down the evidence each one actually produces, run a checker against your map, and find the boundary you left uncovered.

Why Checks Feel Interchangeable Until They Aren't

Most builders start with a mental model that treats validation as a generic safety layer. You sprinkle it on top of the pipeline, like a filter at the end of a drain. Under that model, any check that passes is good news, and more checks are better than fewer.

The stronger model is narrower and more useful: a check is a claim about one boundary, and every boundary proves less than it looks like it does.

A JSON-schema-valid response is parseable and structurally conformant. It is not necessarily true, current, or authorized. A tool call with well-formed arguments matches the tool's contract. It does not mean the side effect is desirable or reversible. A retrieval result that scores high on relevance matched the query. It does not mean the corpus contained the right passage in the first place.

Production incidents live in that space between evidence and guarantee. The check ran. The check passed. The system still failed, because nobody asked what the check was actually capable of establishing.

This exercise makes that question concrete. You will produce a map, run a checker against it, and find the boundary you left empty.

The Five Boundaries You Are Mapping To

A left-to-right request path runs through Input, Retrieval, Model output, Tool execution, and Final verification. Small check markers sit at each boundary; the final marker is visually distinguished to indicate review of the assembled result.
Place each check where it can inspect the relevant evidence; local passes do not replace final verification.

If you have already traced the full request lifecycle — input, context assembly, model call, tools, verification — you have the stage order. Here we only need to fix the boundary names the exercise uses, because consistent labels are what make the mapping gradeable.

BoundaryWhat it coversWhat a check here can establish
InputWhat the user or upstream system supplied, before any model sees itShape, size, encoding, and basic admissibility of the request
RetrievalWhat evidence was selected for this requestRelevance, freshness, and authorization of the selected passages
Model outputWhat the model produced, treated as untrusted dataStructure, schema conformance, and format
Tool executionWhat the system is about to do in the worldArgument validity, permissions, and side-effect scope
Final verificationWhether the assembled result is acceptable to returnEnd-to-end acceptability of the whole response

Here is the distinction the table cannot carry on its own: lifecycle placement decides when a check can act; the inputs it inspects and the rule it applies decide what it can establish. Moving a check later in the pipeline does not widen its view. A schema validator that only ever reads the model's JSON stays a schema validator whether you run it at the model output boundary or at final verification. It still cannot tell you whether retrieval selected the right passage or whether a tool call is safe to execute.

So when you map a check, you are answering two separate questions. Where does it run? And what does it actually look at? A check placed at final verification that inspects only one field is still a local check. A check placed early that inspects the full request trace can still make a broad claim. Position and scope are related, but they are not the same thing — and confusing them is the most common way builders overestimate their coverage.

Knowledge check

Check your understanding

Answer this question before you continue.

A proposed check verifies that a tool call has valid arguments, required permissions, and a constrained side-effect scope. Which boundary is the best fit?
Single Choice

Focus: Map a check about tool permissions and side-effect scope to the lifecycle boundary it can inspect.

Set Up the Exercise Files and Run the Checker

Dependencies are minimal on purpose: Python 3 standard library only. No packages, no API keys, no network calls. The whole exercise runs offline and deterministically.

You need three supplied files, provided alongside this article in the exercise bundle:

  • lifecycle.json — the boundary definitions and stage order
  • scenarios.json — the request scenarios and their proposed checks
  • map_checks.py — the checker

You write one file: answers.json. It is a JSON array — one object per proposed check in scenarios.json. Each object maps the check to a boundary and carries two free-text fields: what evidence the check establishes, and what it cannot prove.

[
  {
    "check_id": "c3",
    "boundary": "model_output",
    "establishes": "response parses and matches the declared schema",
    "cannot_prove": "the values are true, current, or authorized"
  }
]

The snippet above is a fragment, not a complete answer file. Your answers.json needs an entry for every check_id that appears in scenarios.json, or the checker will report the missing ones as unmapped.

Then run the checker:

python map_checks.py lifecycle.json scenarios.json answers.json

Read the output rather than skimming it. Expect two kinds of feedback: per-check pass/fail on boundary placement, and a list of uncovered gaps — scenarios where no check is assigned to a boundary that matters. A representative run looks like this:

check c1  boundary=input            PASS
check c2  boundary=retrieval        PASS
check c3  boundary=model_output     PASS
check c4  boundary=tool_execution   PASS

uncovered gaps:
  scenario s2: final_verification has no assigned check
  scenario s4: retrieval has no assigned check

The checker is deterministic because the mapping is a classification problem. Placement can be graded objectively. The reasoning fields are yours to defend, and that is where the real learning happens.

Reading the Feedback: Evidence Versus Guarantee

A correct boundary placement can still carry a wrong evidence claim. The checker flags placement; the reasoning fields are where you find out whether you actually understand the check.

Here is the pattern to internalize, boundary by boundary:

What this check establishesWhat it cannot prove
Input is well-formed and within limitsThe request is legitimate or well-intentioned
Retrieved passages match the queryThe corpus contained the right passage at all
Output is parseable and schema-conformantThe values are true, current, or authorized
Tool arguments match the tool contractThe side effect is desirable or reversible
The assembled result is acceptable to returnThat it will be acceptable tomorrow, or to a different user

The recurring trap is treating a boundary check as a system guarantee. Green at input, green at retrieval, green at output — so you skip final verification. That is exactly the sequence that ships a confident, well-structured, wrong answer.

Common mistake: Assuming that because every stage passed its own check, the end-to-end result is verified. Each stage check is a local claim. Only a check that actually inspects the assembled result — and is scoped to judge it — can make a claim about the whole.

Knowledge check

Check your understanding

Answer this question before you continue.

A model response parses and passes its declared JSON schema. What is the strongest conclusion this check supports?
Scenario Interpretation

Focus: Distinguish structural evidence from truth, freshness, or authorization guarantees.

Find the Documented Boundary Gap

The scenarios are constructed so at least one boundary has no check assigned. The checker reports it as an uncovered gap, naming the scenario and the boundary — not the fix. Deciding whether the gap is a real risk or an acceptable omission is your job.

Two gap types matter, and they are not the same:

  • A missing check — nothing runs at that boundary.
  • A misplaced check — something runs, but at the wrong boundary, so it proves the wrong thing.

The second is more dangerous because it looks like coverage. A check sitting at the model output boundary that was meant to validate retrieval will pass every time and catch nothing.

Do not assume the gap at final verification is always the serious one. Sometimes the expensive gap is at retrieval, where a bad selection poisons everything downstream — the model reasons correctly over the wrong evidence, and every later check confirms the reasoning while missing the poison.

Not every gap needs a new check. Some need a documented assumption, a human review step, or a narrower scenario scope. The judgment is yours; the checker only tells you where the hole is.

Knowledge check

Check your understanding

Answer this question before you continue.

A check intended to catch a bad retrieval selection is assigned to model output and passes. Which diagnosis best matches the article's warning?
Debugging

Focus: Diagnose why a check assigned to the wrong boundary can appear to provide coverage while missing the intended failure.

Revise the Map: Move a Check, Change a Scenario

Reading the feedback teaches you to interpret. Revising the map teaches you to predict. Do both modifications.

Modification A: Move one check to a different boundary. Before you run the checker, write down what you expect it to report.

Modification B: Alter a scenario so a previously sufficient check becomes insufficient. Observe which gap appears.

Then re-run the same command against the revised answers.json and diff the feedback against the first run.

Watch for these debugging signals:

  • A check that passes placement but fails the evidence field — you classified it right and described it wrong.
  • A scenario that reports two gaps after your edit — your change removed coverage you did not notice.
  • A boundary that now has redundant checks — you added coverage where it already existed and left the real gap open.

This matters beyond the exercise. In a real system, moving a check changes cost, latency, and blast radius. A check at the tool boundary is far more expensive to get wrong than one at the input boundary, because by the time you reach tool execution, the system is about to act in the world.

Knowledge check

Check your understanding

Answer this question before you continue.

After altering a scenario, the checker reports two uncovered gaps. What is the most useful interpretation to investigate first?
Scenario Interpretation

Focus: Interpret checker feedback after a scenario change that removes coverage from more than one boundary.

Turning the Map Into a Habit

The exercise is a stand-in for a decision you make every time you add a check to a real pipeline. The rule is short:

Before adding a check, name the boundary, name the evidence, and name what it cannot prove. If you cannot fill all three, the check is decoration.

A few heuristics that follow from the mapping:

  • Put the cheapest check that can still catch the failure as early as possible — but never let an early check stand in for final verification.
  • Watch for check sprawl: many overlapping checks that all prove the same thing while one boundary stays empty.
  • Sometimes the right move is not a check at all. Narrow the scope, add human review, or refuse the request outright.

I have watched builders add a fifth validator to a boundary that already had three, while retrieval ran with no check whatsoever. The instinct to add checks is not the same as the judgment to place them.

Where to Go Next

Take one real request from a pipeline you already own. Map it against the same five boundaries. Count how many have no check at all — not a weak check, not a misplaced one, but nothing.

That count is the gap you have been shipping with. The exercise gave you the vocabulary and the checker. Your own pipeline gives you the stakes.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

A schema validator reads only the model's JSON. If it is moved from the model-output boundary to final verification, what new conclusion can it support solely because of that move?
Question 1 of 2Comparison Reasoning

Focus: Explain why moving a check later in the lifecycle does not expand the evidence it inspects.

Before adding a proposed check to a request pipeline, which description best follows the article's recommended habit?
Question 2 of 2Single Choice

Focus: Apply the article's three-part habit for evaluating a proposed check.

References

  1. Evaluation concepts - Docs by LangChaindocs.langchain.com
  2. How to evaluate LLMs for SQL generationdevelopers.openai.com
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.