Practice Applying Agent Stopping Rules to Loop Traces
You can recite the four dispositions — continue, stop, retry, handoff — and still stare at a raw trace with no idea which one applies. That gap is the…

Key topics
Six events. One decision. Most learners freeze right there.
You can recite the four dispositions — continue, stop, retry, handoff — and still stare at a raw trace with no idea which one applies. That gap is the whole problem. Stopping rules are not a setting you configure once and forget. They are a judgment you make, step by step, from evidence in the trace. This exercise trains that judgment until it becomes automatic.
The prerequisite article covered designing budgets and stopping rules. This one assumes you have that mental model and drills the reading skill on top of it: given a trace and a policy, commit to a bounded, evidence-supported disposition.
Why Trace Reading Is the Skill, Not the Theory
A trace is a sequence of state changes. That is the entire frame you need.
Progress means the state moved toward the goal — new information, a new state, a narrower search space. Repetition means the state returned to something already seen — same action, same inputs, same or worse result. Every disposition decision flows from that first call.
The common failure looks like this: a learner reads the last event, pattern-matches on its shape, and picks a disposition. That is guessing with extra steps. A disposition without a cited event is a guess, and guesses do not survive a policy change. When someone tightens a budget or moves a handoff threshold, the guesser has to start over. The reader who can point at event 3 and event 5 just re-applies the rule.
Budgets bound the loop. This exercise trains the decision inside the bound.
Set Up the Exercise Files
Three files, Python 3 standard library only. No installs, no API keys, no network calls.
traces.json— the cases. Each trace is an ordered list of events. Every event carries an action label and an observable result or state field.policy.json— the rules that decide. Thresholds and conditions that map a trace pattern to a disposition.check_traces.py— the verifier. It compares your answers against the policy.
Your job produces answers.json: one entry per trace, containing a progress/repetition label, a disposition, and the cited event indices.
A trace entry looks roughly like this:
{
"trace_id": "t1",
"events": [
{"index": 0, "action": "search", "result": "no_match"},
{"index": 1, "action": "search", "result": "no_match"},
{"index": 2, "action": "refine_query", "result": "partial_match"}
]
}
Your answer entry mirrors it:
{
"trace_id": "t1",
"label": "progress",
"disposition": "continue",
"evidence": [2]
}
Before you commit to real answers, run the checker against an empty or partial answers file. See the feedback format first. You want to know what a mismatch looks like before you are emotionally invested in being right.
python check_traces.py traces.json policy.json answers.json
Read One Trace Before You Decide Anything
Separate reading from deciding. Three passes, in order.
Pass one: walk the events in sequence. Mark where the state changed and where it did not. Do not consult the policy yet.
Pass two: compare each event against earlier events. Look for a repeated action with the same or worse result. This is where repetition becomes visible — not in any single event, but in the relationship between events.
Pass three: only now open policy.json and map the pattern you found to a disposition.
Why the separation matters: the most common error is pattern-matching on the last event instead of the sequence. A trace that ends with a successful search looks like progress until you notice the same search failed twice before it. The sequence is the evidence. The last event is just the last event.
Progress or Repetition: The First Call You Make
This label is load-bearing. Get it wrong and the disposition is wrong even if you apply the policy perfectly.
| Test | Question | Signal |
|---|---|---|
| Progress | Did this produce new information, a new state, or a narrower search space? | State moved toward the goal |
| Repetition | Same action, same inputs, same or worse result? | Budget spent, nothing moved |
The ambiguous middle is where people stumble: a retry with changed inputs is not repetition. A retry with identical inputs is. The action label is identical in both cases. The arguments are what changed.
Common mistake: Treating any repeated action label as repetition when the arguments actually changed.
search("cats")followed bysearch("cats and dogs")is two different actions wearing the same label. Read the arguments, not the label.
Knowledge check
Check your understanding
Answer this question before you continue.
Choosing Continue, Stop, Retry, or Handoff
Four dispositions, each tied to trace evidence and a policy field.
Continue — progress is visible and the budget is not exhausted. The loop is working. Let it work.
Stop — the goal condition is met, or the budget is exhausted with no path forward. Both are terminal, but they mean different things. One is success. One is a bounded failure.
Retry — repetition is detected but the failure looks recoverable. The policy must say what changes on the retry. A retry that changes nothing is just a slower loop.
Handoff — the trace shows ambiguity, risk, or a decision the policy explicitly routes to a human.
The distinction that trips people up: retry fixes a recoverable failure; handoff escalates an unresolvable one. Choosing retry when the policy requires handoff is the classic mistake. It feels productive. It burns budget on a problem the loop cannot solve.
When you cite evidence, name the event indices. Not "the trace shows repeated searches" — [0, 1]. Specific indices survive a policy change. General shapes do not.
Knowledge check
Check your understanding
Answer this question before you continue.
Run the Checker and Read the Feedback
python check_traces.py traces.json policy.json answers.json
The checker verifies your disposition and cited evidence against the policy. It does not judge your prose reasoning. That boundary matters: a correct disposition with a wrong citation is a lucky guess, and the checker will tell you so.
When you get a mismatch, isolate which layer failed:
- Label wrong? You misread progress as repetition, or the reverse. Go back to pass two.
- Disposition wrong? The label was right but you mapped it to the wrong policy branch. Re-read the policy field.
- Citation wrong? You reached the right answer for the wrong reason. This is the most dangerous kind of correct.
Note: A wrong answer with the right citation is more useful than a right answer with no citation. The first teaches you where your rule broke. The second teaches you nothing.
Knowledge check
Check your understanding
Answer this question before you continue.
Malformed Traces and Unknown Actions Are Signals
Broken input is diagnostic evidence, not a crash to work around.
Malformed event sequences — out-of-order events, missing result fields, or a trace that ends without a terminal state. These tell you something about the harness that produced the trace.
Unknown action labels — the policy has no rule for them. The correct disposition is not "guess the nearest known label." It is to flag the gap.
Why this matters: an unhandled action label means the policy is incomplete. That is a finding worth recording, not a bug to paper over. Silently mapping an unknown label to the closest known one hides the gap and corrupts every downstream decision.
Warning: If you find yourself inventing a rule the policy does not contain, stop. You are no longer applying the policy. You are writing a new one, and nobody reviewed it.
Knowledge check
Check your understanding
Answer this question before you continue.
Change One Thing and Decide Again
Now prove your decision rule is evidence-based rather than memorized.
Experiment A: Change one event in a trace. Flip a repeated action into a productive one. Re-run the checker. Does your disposition change? Can you say which event caused the change?
Experiment B: Change one policy field. Tighten a budget or move a handoff threshold. Re-decide the same trace. Does the disposition shift in the direction the policy change predicts?
A stable decision rule changes only when the evidence or the policy changes — and you can name which. An unstable rule flips for reasons you cannot articulate. That instability is the tell: you were pattern-matching, not deciding.
Record both runs. Write down the before and after. Memory will lie to you about what you originally answered.
Where This Judgment Pays Off
The same reading pass applies to production traces, logs, and observability output from real agent runs. The traces are longer and noisier. The policy is often implicit until you write it down. But the sequence — label progress or repetition, cite the event, choose the disposition the policy supports — does not change.
Bounded, evidence-supported dispositions are what make an agent loop reviewable by someone other than its author. When you can point at event 3 and event 5, a teammate can check your reasoning. When you cannot, they have to trust you. Trust does not scale. Evidence does.
The habit worth keeping: cite the event before you name the disposition.
Your Next Move
Take one real trace from a workflow you already run. Write down the policy you are implicitly using — the thresholds, the retry conditions, the handoff triggers that live in your head. Then check whether your own decisions would survive the same perturbation test: change one event, change one threshold, and see if your disposition still holds.
If it does, you have a rule. If it does not, you have a habit — and now you know the difference.
Knowledge check
Final check
Finish the article by checking the ideas you just learned.
References
Want a more structured LLMOps path?
Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.
Large Language Models Starter Pack
A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.
- 227-page Illustrated PDF edition
- 12 guided LLM engineering chapters
- Visual concept diagrams
- Self-assessment quizzes
- Bonus deep-dive sections
- Prompt design, structured output, context windows & RAG pipelines
- Agents, tool calling, prompt injection, evaluation & application lifecycles
Coming soon


