Skip to content
intermediate

Practice Triage: Find the Likely Cause of an LLM Agent Failure

The run failed. The trace is four hundred lines long. Your instinct says the model is dumb — and that instinct is almost always the first thing to throw…

Published 2026-10-03Updated 2026-10-049 min read
Detailed texture of golden sand at Broadstairs Beach, England, showcasing nature's artistry.
Detailed texture of golden sand at Broadstairs Beach, England, showcasing nature's artistry. Photo by Kate Strilchuk on Pexels.

The run failed. The trace is four hundred lines long. Your instinct says the model is dumb — and that instinct is almost always the first thing to throw away.

Here is the discipline I want you to practice: name the failure class, cite the field that proves it, and pick the smallest next check that would change your mind. That last part matters more than the label. A diagnosis you cannot test is just a story you told yourself about a log file.

This drill gives you synthetic incident records and a checker. You classify each case, record evidence, choose a next diagnostic step, then run the script and see what it says. Then you break one of your own answers on purpose and watch what happens to your confidence.

Why Agent Failures Resist a Single Explanation

If you have already worked through the common agent failure modes, you know the five buckets: planning, tool selection, execution, state, and stopping. What that article gives you is the taxonomy. What it does not give you is reps.

The gap between knowing the buckets and using them shows up the moment a real trace lands in front of you. A loop looks like a stopping problem. A wrong answer looks like a planning problem. A stalled run looks like a tool problem. Symptoms are loud and ambiguous; causes are quiet and specific. One symptom can map to three different classes, and the fix for the wrong class is worse than no fix at all — it hides the symptom while the cause keeps charging rent.

Triage is a classification problem with a discipline attached. Classify first. Then choose the smallest diagnostic step that would confirm or falsify your class. Fixes come later, after the class is settled.

Set Up the Triage Drill

The drill runs on the Python 3 standard library. No installs, no API keys, no network calls. Everything is local, synthetic, and safe to break.

You need three files:

  • triage_cases.py — the checker
  • incidents.json — synthetic incident records
  • answers.json — your classifications, evidence, and next steps

Run it like this:

python3 triage_cases.py incidents.json answers.json

The checker reads your answers, compares them against an answer key, and prints per-case feedback. You get a verdict per case — correct, incorrect, or uncertain — plus a summary count at the end.

Note: The records are synthetic and self-contained. That means results are reproducible, and a wrong answer costs you nothing but a re-read. Break things freely.

Read One Incident Record Before You Classify

Before you touch answers.json, read one record end to end. Each incident has the same shape:

  • task — what the agent was asked to do
  • plan steps — the decomposition the agent produced
  • tool calls — which tools were invoked, with what arguments
  • tool results — what came back
  • state snapshots — what the agent carried forward between steps
  • stop condition — the rule that was supposed to end the run
  • final output — what the agent actually returned

Some of these fields are direct evidence. Some are only suggestive. The tool call log tells you what happened. The plan tells you what the agent intended. The state snapshots tell you what survived.

Here is the trap. Read only the final output and you will classify from the symptom. Read the plan, the calls, and the state together and the same record can support two different classes. That ambiguity is not a flaw in the exercise — it is the exercise. Your job is to find the field that separates the two readings, or to admit that no field does.

Establish the habit now: quote the field, not your impression of the field.

Knowledge check

Check your understanding

Answer this question before you continue.

You need to verify which tool the agent actually invoked and the arguments it supplied. Which field should you inspect first?
Single Choice

Focus: Select the incident field that directly records which tools were invoked and with what arguments.

The Five Classes and Their Evidence Signatures

Classification stops being a vibe check when each class has an observable signature. Here is what to look for.

ClassSignatureWhat stays correct
PlanningDecomposition is wrong or missing a stepDownstream steps execute as designed
Tool selectionRight capability exists, wrong tool chosen — or no tool chosen when one was neededThe plan itself is sound
ExecutionCorrect tool called with bad arguments, or the call erroredThe model's intent was sound
StateInformation lost, overwritten, or not carried forwardPlan and tools were fine
StoppingAgent never recognized completion, or stopped earlyThe work itself was correct
Insufficient evidenceRecord cannot distinguish two classes—

The right-hand column is the part beginners skip. A planning failure leaves the individual steps intact. An execution failure leaves the intent intact. A state failure leaves the plan and the tools intact. When you can name what still works, you have narrowed the class.

Insufficient evidence is a real answer. If two classes fit the record equally well and no field separates them, saying so is correct. Over-claiming on a thin record is the most common way to get this drill wrong.

Knowledge check

Check your understanding

Answer this question before you continue.

An agent's decomposition omits a required step. The steps it did plan are carried out as designed, with no tool errors or lost state reported. Which class best fits the evidence?
Scenario Interpretation

Focus: Classify a failure as planning when the decomposition is wrong but downstream steps execute as designed.

Choose the Smallest Justified Next Step

A flowchart moves from a symptom to two candidate failure classes, then to a field that can separate them. Different field outcomes support either class; if no field separates them, the result is insufficient evidence.
A useful next check separates competing explanations; if the record cannot do that, keep the case uncertain.

Classifying is half the job. The other half is deciding what to check next — and here the bar is higher than most people expect.

A diagnostic step is justified when it would change your classification if it came back the other way. If both outcomes leave you with the same answer, you are not diagnosing. You are confirming.

Prefer the cheapest check that separates your top two candidate classes. One log field. One replayed step. One re-run with a fixed input. That is usually enough.

Common mistake: Proposing a fix instead of a diagnostic. "Add a retry" is not a diagnostic step. "Check whether the tool result was written to state before step four" is.

Reject any step that only confirms what you already believe. The whole point of triage is to find the check that could embarrass you.

Knowledge check

Check your understanding

Answer this question before you continue.

A later step behaves as if an earlier tool result is missing. You are considering a state failure, but want a check that could change that classification. Which next step best meets that goal?
Comparison Reasoning

Focus: Choose a low-cost diagnostic check that can distinguish a state failure from another candidate cause.

Record Your Evidence and Run the Checker

Each entry in answers.json needs three things: a class, a supporting evidence field, and a next diagnostic step. The evidence field is where most people get lazy — they write a plausible story instead of citing the field that supports the class.

Run the script and read the feedback carefully. The checker distinguishes between a wrong class and a right class with weak evidence. Those are different failures, and they need different corrections.

A correct class with unsupported evidence is still a weak answer. The checker should push you back toward the field you skipped.

Treat the summary counts as a baseline report, not a score to optimize. You are measuring the quality of your reasoning, not your ability to match a key.

Knowledge check

Check your understanding

Answer this question before you continue.

The checker says your class is correct but your supporting evidence is weak. What should you do next?
Debugging

Focus: Use checker feedback to distinguish an unsupported evidence field from an incorrect class.

Edit One Evidence Field and Decide Again

Now the experiment that makes this drill worth running.

Pick a case you classified confidently. Find the single field your answer depended on. Remove it, or change it, and re-run.

Then re-decide. Does the class stay the same? Does it flip? Or does it collapse into insufficient evidence?

This is the real lesson. A diagnosis is only as strong as the evidence it rests on, and most confident diagnoses are one deleted field away from uncertainty. Some classes are robust to missing evidence — a clear execution error usually stays an execution error even without the state snapshots. Others are fragile. State failures in particular tend to dissolve the moment you lose the snapshot that showed the loss.

Note which classes held and which fell apart. That pattern tells you where your instrumentation is thin.

Where Triage Habits Break Down

The triage process has its own failure modes. Watch for these.

Confirmation bias. You guess a class on the first read, then read the trace looking for support. The fix is to write down two candidate classes before you commit to one.

Symptom anchoring. A loop looks like a stopping problem. Often it is a state problem — the agent loops because it never received the information that would let it stop. Classify the cause, not the shape of the failure.

Over-claiming on thin records. Beginners hate writing "insufficient evidence." It feels like a non-answer. It is not. It is the honest answer when the record cannot decide, and it points directly at the field you need to add.

Skipping the evidence field. A plausible story is not a citation. If you cannot name the field, you do not have a diagnosis yet.

Extend the Drill to Your Own Traces

The synthetic records are training wheels. The real practice starts when you apply the same three fields — class, evidence, smallest next step — to a trace from your own agent.

Do this once this week. Take a recent failure, write the three fields, and see how far you get before you have to admit the record does not contain what you need.

Then add one observability field that would have made that failure easier to classify. A state snapshot at each step. A log line when the stop condition is evaluated. One field, chosen because a real failure demanded it.

Keep a running list of ambiguous cases. Those are not evidence that your model is weak. They are evidence that your instrumentation is thin — and that is a much cheaper thing to fix.

The Rule to Carry

Classify before you fix. Cite the field before you claim. Mark insufficient evidence when the record cannot decide.

That is the whole habit. It will not make failures disappear, but it will stop you from spending an afternoon patching a symptom while the cause keeps running.

Your next step: instrument one real agent run with the two fields that made these drill cases classifiable — a state snapshot per step and a stop-condition log. Then re-run the triage on your own trace and see which class your failures actually belong to.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

A trace shows an agent looping, but contains neither state snapshots nor a log of stop-condition evaluation. The loop could result from missing information or failure to recognize completion. What is the most justified classification from this record?
Question 1 of 2Scenario Interpretation

Focus: Avoid inferring a failure class from a loop symptom when the record cannot distinguish state from stopping.

A case shows a correct tool call that returned an explicit error. You remove the state snapshots, but the call and error remain visible. Which conclusion is best supported?
Question 2 of 2Comparison Reasoning

Focus: Reassess a diagnosis after removing evidence and retain it when independent direct evidence remains.

Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.

Close-up of a car dashboard at night with illuminated speedometer and tech displays.
intermediate
10 min read

Build a Simple LLM Agent

Most beginners expect an agent to be a special kind of model—something with built-in magic that can browse the web, run code, and get things done. Then…

Read tutorial