Skip to content
intermediate

Practice Applying Agent Stopping Rules to Loop Traces

You can recite the four dispositions — continue, stop, retry, handoff — and still stare at a raw trace with no idea which one applies. That gap is the…

Published 2026-10-03Updated 2026-10-048 min read
Vibrant close-up of network cable connectors with colorful lighting.
Vibrant close-up of network cable connectors with colorful lighting. Photo by Nic Wood on Pexels.

Six events. One decision. Most learners freeze right there.

You can recite the four dispositions — continue, stop, retry, handoff — and still stare at a raw trace with no idea which one applies. That gap is the whole problem. Stopping rules are not a setting you configure once and forget. They are a judgment you make, step by step, from evidence in the trace. This exercise trains that judgment until it becomes automatic.

The prerequisite article covered designing budgets and stopping rules. This one assumes you have that mental model and drills the reading skill on top of it: given a trace and a policy, commit to a bounded, evidence-supported disposition.

Why Trace Reading Is the Skill, Not the Theory

A trace is a sequence of state changes. That is the entire frame you need.

Progress means the state moved toward the goal — new information, a new state, a narrower search space. Repetition means the state returned to something already seen — same action, same inputs, same or worse result. Every disposition decision flows from that first call.

The common failure looks like this: a learner reads the last event, pattern-matches on its shape, and picks a disposition. That is guessing with extra steps. A disposition without a cited event is a guess, and guesses do not survive a policy change. When someone tightens a budget or moves a handoff threshold, the guesser has to start over. The reader who can point at event 3 and event 5 just re-applies the rule.

Budgets bound the loop. This exercise trains the decision inside the bound.

Set Up the Exercise Files

Three files, Python 3 standard library only. No installs, no API keys, no network calls.

  • traces.json — the cases. Each trace is an ordered list of events. Every event carries an action label and an observable result or state field.
  • policy.json — the rules that decide. Thresholds and conditions that map a trace pattern to a disposition.
  • check_traces.py — the verifier. It compares your answers against the policy.

Your job produces answers.json: one entry per trace, containing a progress/repetition label, a disposition, and the cited event indices.

A trace entry looks roughly like this:

{
  "trace_id": "t1",
  "events": [
    {"index": 0, "action": "search", "result": "no_match"},
    {"index": 1, "action": "search", "result": "no_match"},
    {"index": 2, "action": "refine_query", "result": "partial_match"}
  ]
}

Your answer entry mirrors it:

{
  "trace_id": "t1",
  "label": "progress",
  "disposition": "continue",
  "evidence": [2]
}

Before you commit to real answers, run the checker against an empty or partial answers file. See the feedback format first. You want to know what a mismatch looks like before you are emotionally invested in being right.

python check_traces.py traces.json policy.json answers.json

Read One Trace Before You Decide Anything

Separate reading from deciding. Three passes, in order.

Pass one: walk the events in sequence. Mark where the state changed and where it did not. Do not consult the policy yet.

Pass two: compare each event against earlier events. Look for a repeated action with the same or worse result. This is where repetition becomes visible — not in any single event, but in the relationship between events.

Pass three: only now open policy.json and map the pattern you found to a disposition.

Why the separation matters: the most common error is pattern-matching on the last event instead of the sequence. A trace that ends with a successful search looks like progress until you notice the same search failed twice before it. The sequence is the evidence. The last event is just the last event.

Progress or Repetition: The First Call You Make

This label is load-bearing. Get it wrong and the disposition is wrong even if you apply the policy perfectly.

TestQuestionSignal
ProgressDid this produce new information, a new state, or a narrower search space?State moved toward the goal
RepetitionSame action, same inputs, same or worse result?Budget spent, nothing moved

The ambiguous middle is where people stumble: a retry with changed inputs is not repetition. A retry with identical inputs is. The action label is identical in both cases. The arguments are what changed.

Common mistake: Treating any repeated action label as repetition when the arguments actually changed. search("cats") followed by search("cats and dogs") is two different actions wearing the same label. Read the arguments, not the label.

Knowledge check

Check your understanding

Answer this question before you continue.

A trace contains `search("cats")` followed by `search("cats and dogs")`. What does the article say to check before labeling this repetition?
Scenario Interpretation

Focus: Classify a repeated action label by comparing its inputs and results with earlier events.

Choosing Continue, Stop, Retry, or Handoff

A flowchart compares event inputs and results, branches to progress or repetition, then routes through a policy check to continue, stop, retry, or handoff.
Compare the events first; let the policy—not the last event alone—determine the next action.

Four dispositions, each tied to trace evidence and a policy field.

Continue — progress is visible and the budget is not exhausted. The loop is working. Let it work.

Stop — the goal condition is met, or the budget is exhausted with no path forward. Both are terminal, but they mean different things. One is success. One is a bounded failure.

Retry — repetition is detected but the failure looks recoverable. The policy must say what changes on the retry. A retry that changes nothing is just a slower loop.

Handoff — the trace shows ambiguity, risk, or a decision the policy explicitly routes to a human.

The distinction that trips people up: retry fixes a recoverable failure; handoff escalates an unresolvable one. Choosing retry when the policy requires handoff is the classic mistake. It feels productive. It burns budget on a problem the loop cannot solve.

When you cite evidence, name the event indices. Not "the trace shows repeated searches" — [0, 1]. Specific indices survive a policy change. General shapes do not.

Knowledge check

Check your understanding

Answer this question before you continue.

A trace shows ambiguity that the policy explicitly routes to a human. Which disposition follows the article's distinction?
Single Choice

Focus: Choose handoff rather than retry when trace evidence meets a policy-defined escalation condition.

Run the Checker and Read the Feedback

python check_traces.py traces.json policy.json answers.json

The checker verifies your disposition and cited evidence against the policy. It does not judge your prose reasoning. That boundary matters: a correct disposition with a wrong citation is a lucky guess, and the checker will tell you so.

When you get a mismatch, isolate which layer failed:

  • Label wrong? You misread progress as repetition, or the reverse. Go back to pass two.
  • Disposition wrong? The label was right but you mapped it to the wrong policy branch. Re-read the policy field.
  • Citation wrong? You reached the right answer for the wrong reason. This is the most dangerous kind of correct.

Note: A wrong answer with the right citation is more useful than a right answer with no citation. The first teaches you where your rule broke. The second teaches you nothing.

Knowledge check

Check your understanding

Answer this question before you continue.

The checker reports a mismatch because your disposition is correct but your cited event indices are wrong. Which layer should you investigate?
Debugging

Focus: Use checker feedback to distinguish a correct disposition from evidence-supported reasoning.

Malformed Traces and Unknown Actions Are Signals

Broken input is diagnostic evidence, not a crash to work around.

Malformed event sequences — out-of-order events, missing result fields, or a trace that ends without a terminal state. These tell you something about the harness that produced the trace.

Unknown action labels — the policy has no rule for them. The correct disposition is not "guess the nearest known label." It is to flag the gap.

Why this matters: an unhandled action label means the policy is incomplete. That is a finding worth recording, not a bug to paper over. Silently mapping an unknown label to the closest known one hides the gap and corrupts every downstream decision.

Warning: If you find yourself inventing a rule the policy does not contain, stop. You are no longer applying the policy. You are writing a new one, and nobody reviewed it.

Knowledge check

Check your understanding

Answer this question before you continue.

A trace contains an action label for which the policy has no rule. What is the article's recommended response?
Misconception Check

Focus: Respond to an action label absent from the policy without inventing a rule.

Change One Thing and Decide Again

Now prove your decision rule is evidence-based rather than memorized.

Experiment A: Change one event in a trace. Flip a repeated action into a productive one. Re-run the checker. Does your disposition change? Can you say which event caused the change?

Experiment B: Change one policy field. Tighten a budget or move a handoff threshold. Re-decide the same trace. Does the disposition shift in the direction the policy change predicts?

A stable decision rule changes only when the evidence or the policy changes — and you can name which. An unstable rule flips for reasons you cannot articulate. That instability is the tell: you were pattern-matching, not deciding.

Record both runs. Write down the before and after. Memory will lie to you about what you originally answered.

Where This Judgment Pays Off

The same reading pass applies to production traces, logs, and observability output from real agent runs. The traces are longer and noisier. The policy is often implicit until you write it down. But the sequence — label progress or repetition, cite the event, choose the disposition the policy supports — does not change.

Bounded, evidence-supported dispositions are what make an agent loop reviewable by someone other than its author. When you can point at event 3 and event 5, a teammate can check your reasoning. When you cannot, they have to trust you. Trust does not scale. Evidence does.

The habit worth keeping: cite the event before you name the disposition.

Your Next Move

Take one real trace from a workflow you already run. Write down the policy you are implicitly using — the thresholds, the retry conditions, the handoff triggers that live in your head. Then check whether your own decisions would survive the same perturbation test: change one event, change one threshold, and see if your disposition still holds.

If it does, you have a rule. If it does not, you have a habit — and now you know the difference.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

You change one event in a trace, rerun the checker, and the disposition changes. Which follow-up best tests whether the decision was evidence-based?
Question 1 of 2Comparison Reasoning

Focus: Interpret how a bounded decision rule should respond when evidence or policy changes.

A learner sees a promising final event and immediately chooses continue, without comparing it with earlier events. What should they do next to follow the article's method?
Question 2 of 2Scenario Interpretation

Focus: Apply the article's evidence-first sequence before mapping a trace to a policy disposition.

References

  1. Running agents | OpenAI APIdevelopers.openai.com
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.

Close-up of a car dashboard at night with illuminated speedometer and tech displays.
intermediate
10 min read

Build a Simple LLM Agent

Most beginners expect an agent to be a special kind of model—something with built-in magic that can browse the web, run code, and get things done. Then…

Read tutorial