Practice Triage: Find the Likely Cause of an LLM Agent Failure
The run failed. The trace is four hundred lines long. Your instinct says the model is dumb — and that instinct is almost always the first thing to throw…

Key topics
The run failed. The trace is four hundred lines long. Your instinct says the model is dumb — and that instinct is almost always the first thing to throw away.
Here is the discipline I want you to practice: name the failure class, cite the field that proves it, and pick the smallest next check that would change your mind. That last part matters more than the label. A diagnosis you cannot test is just a story you told yourself about a log file.
This drill gives you synthetic incident records and a checker. You classify each case, record evidence, choose a next diagnostic step, then run the script and see what it says. Then you break one of your own answers on purpose and watch what happens to your confidence.
Why Agent Failures Resist a Single Explanation
If you have already worked through the common agent failure modes, you know the five buckets: planning, tool selection, execution, state, and stopping. What that article gives you is the taxonomy. What it does not give you is reps.
The gap between knowing the buckets and using them shows up the moment a real trace lands in front of you. A loop looks like a stopping problem. A wrong answer looks like a planning problem. A stalled run looks like a tool problem. Symptoms are loud and ambiguous; causes are quiet and specific. One symptom can map to three different classes, and the fix for the wrong class is worse than no fix at all — it hides the symptom while the cause keeps charging rent.
Triage is a classification problem with a discipline attached. Classify first. Then choose the smallest diagnostic step that would confirm or falsify your class. Fixes come later, after the class is settled.
Set Up the Triage Drill
The drill runs on the Python 3 standard library. No installs, no API keys, no network calls. Everything is local, synthetic, and safe to break.
You need three files:
triage_cases.py— the checkerincidents.json— synthetic incident recordsanswers.json— your classifications, evidence, and next steps
Run it like this:
python3 triage_cases.py incidents.json answers.json
The checker reads your answers, compares them against an answer key, and prints per-case feedback. You get a verdict per case — correct, incorrect, or uncertain — plus a summary count at the end.
Note: The records are synthetic and self-contained. That means results are reproducible, and a wrong answer costs you nothing but a re-read. Break things freely.
Read One Incident Record Before You Classify
Before you touch answers.json, read one record end to end. Each incident has the same shape:
- task — what the agent was asked to do
- plan steps — the decomposition the agent produced
- tool calls — which tools were invoked, with what arguments
- tool results — what came back
- state snapshots — what the agent carried forward between steps
- stop condition — the rule that was supposed to end the run
- final output — what the agent actually returned
Some of these fields are direct evidence. Some are only suggestive. The tool call log tells you what happened. The plan tells you what the agent intended. The state snapshots tell you what survived.
Here is the trap. Read only the final output and you will classify from the symptom. Read the plan, the calls, and the state together and the same record can support two different classes. That ambiguity is not a flaw in the exercise — it is the exercise. Your job is to find the field that separates the two readings, or to admit that no field does.
Establish the habit now: quote the field, not your impression of the field.
Knowledge check
Check your understanding
Answer this question before you continue.
The Five Classes and Their Evidence Signatures
Classification stops being a vibe check when each class has an observable signature. Here is what to look for.
| Class | Signature | What stays correct |
|---|---|---|
| Planning | Decomposition is wrong or missing a step | Downstream steps execute as designed |
| Tool selection | Right capability exists, wrong tool chosen — or no tool chosen when one was needed | The plan itself is sound |
| Execution | Correct tool called with bad arguments, or the call errored | The model's intent was sound |
| State | Information lost, overwritten, or not carried forward | Plan and tools were fine |
| Stopping | Agent never recognized completion, or stopped early | The work itself was correct |
| Insufficient evidence | Record cannot distinguish two classes | — |
The right-hand column is the part beginners skip. A planning failure leaves the individual steps intact. An execution failure leaves the intent intact. A state failure leaves the plan and the tools intact. When you can name what still works, you have narrowed the class.
Insufficient evidence is a real answer. If two classes fit the record equally well and no field separates them, saying so is correct. Over-claiming on a thin record is the most common way to get this drill wrong.
Knowledge check
Check your understanding
Answer this question before you continue.
Choose the Smallest Justified Next Step
Classifying is half the job. The other half is deciding what to check next — and here the bar is higher than most people expect.
A diagnostic step is justified when it would change your classification if it came back the other way. If both outcomes leave you with the same answer, you are not diagnosing. You are confirming.
Prefer the cheapest check that separates your top two candidate classes. One log field. One replayed step. One re-run with a fixed input. That is usually enough.
Common mistake: Proposing a fix instead of a diagnostic. "Add a retry" is not a diagnostic step. "Check whether the tool result was written to state before step four" is.
Reject any step that only confirms what you already believe. The whole point of triage is to find the check that could embarrass you.
Knowledge check
Check your understanding
Answer this question before you continue.
Record Your Evidence and Run the Checker
Each entry in answers.json needs three things: a class, a supporting evidence field, and a next diagnostic step. The evidence field is where most people get lazy — they write a plausible story instead of citing the field that supports the class.
Run the script and read the feedback carefully. The checker distinguishes between a wrong class and a right class with weak evidence. Those are different failures, and they need different corrections.
A correct class with unsupported evidence is still a weak answer. The checker should push you back toward the field you skipped.
Treat the summary counts as a baseline report, not a score to optimize. You are measuring the quality of your reasoning, not your ability to match a key.
Knowledge check
Check your understanding
Answer this question before you continue.
Edit One Evidence Field and Decide Again
Now the experiment that makes this drill worth running.
Pick a case you classified confidently. Find the single field your answer depended on. Remove it, or change it, and re-run.
Then re-decide. Does the class stay the same? Does it flip? Or does it collapse into insufficient evidence?
This is the real lesson. A diagnosis is only as strong as the evidence it rests on, and most confident diagnoses are one deleted field away from uncertainty. Some classes are robust to missing evidence — a clear execution error usually stays an execution error even without the state snapshots. Others are fragile. State failures in particular tend to dissolve the moment you lose the snapshot that showed the loss.
Note which classes held and which fell apart. That pattern tells you where your instrumentation is thin.
Where Triage Habits Break Down
The triage process has its own failure modes. Watch for these.
Confirmation bias. You guess a class on the first read, then read the trace looking for support. The fix is to write down two candidate classes before you commit to one.
Symptom anchoring. A loop looks like a stopping problem. Often it is a state problem — the agent loops because it never received the information that would let it stop. Classify the cause, not the shape of the failure.
Over-claiming on thin records. Beginners hate writing "insufficient evidence." It feels like a non-answer. It is not. It is the honest answer when the record cannot decide, and it points directly at the field you need to add.
Skipping the evidence field. A plausible story is not a citation. If you cannot name the field, you do not have a diagnosis yet.
Extend the Drill to Your Own Traces
The synthetic records are training wheels. The real practice starts when you apply the same three fields — class, evidence, smallest next step — to a trace from your own agent.
Do this once this week. Take a recent failure, write the three fields, and see how far you get before you have to admit the record does not contain what you need.
Then add one observability field that would have made that failure easier to classify. A state snapshot at each step. A log line when the stop condition is evaluated. One field, chosen because a real failure demanded it.
Keep a running list of ambiguous cases. Those are not evidence that your model is weak. They are evidence that your instrumentation is thin — and that is a much cheaper thing to fix.
The Rule to Carry
Classify before you fix. Cite the field before you claim. Mark insufficient evidence when the record cannot decide.
That is the whole habit. It will not make failures disappear, but it will stop you from spending an afternoon patching a symptom while the cause keeps running.
Your next step: instrument one real agent run with the two fields that made these drill cases classifiable — a state snapshot per step and a stop-condition log. Then re-run the triage on your own trace and see which class your failures actually belong to.
Knowledge check
Final check
Finish the article by checking the ideas you just learned.
Want a more structured LLMOps path?
Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.
Large Language Models Starter Pack
A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.
- 227-page Illustrated PDF edition
- 12 guided LLM engineering chapters
- Visual concept diagrams
- Self-assessment quizzes
- Bonus deep-dive sections
- Prompt design, structured output, context windows & RAG pipelines
- Agents, tool calling, prompt injection, evaluation & application lifecycles
Coming soon


