Skip to content
intermediate

LLM Tool-Calling Control Flow: States, Results, and Failure Boundaries

That gap is not bad luck. It is a missing model. Most beginners picture tool calling as a straight line: ask, call, answer. That picture is accurate for…

Published 2026-10-03Updated 2026-10-048 min read
Close-up of a video editing timeline on a computer screen, showcasing modern technology.
Close-up of a video editing timeline on a computer screen, showcasing modern technology. Photo by Vito Goričan on Pexels.

The demo worked. The production run charged the card twice.

That gap is not bad luck. It is a missing model. Most beginners picture tool calling as a straight line: ask, call, answer. That picture is accurate for exactly one case — a single call that cannot fail, cannot change the world, and never needs repeating. The moment any of those three conditions breaks, the straight line becomes a loop over mutable state, and the loop has failure boundaries you cannot see until something goes wrong.

The useful question is not "did the call succeed?" It is "which state transition did this attempt, and is that transition safe to repeat?" Answer that, and the duplicate charge, the stale result, and the infinite loop stop being mysteries. They become locations.

Why the Straight-Line Model Breaks

Grant the straight-line model its narrow win. For a read-only lookup that always succeeds, request → tool call → answer is correct, and building anything more is overhead.

It breaks under three conditions:

  • The call can fail. Networks time out, APIs throw, arguments arrive malformed.
  • The call changes the world. A write, a send, a charge, a delete — anything that survives a retry.
  • The call must be repeated. Retries, loops, and multi-step tasks all require running the same intent more than once.

Here is the hidden constraint that the straight line hides: the model does not hold state. Your application does. The model proposes. Your code decides what happened, what to record, and what to do next. Every real decision about state lives outside the model.

So before the detail, here is the map. One tool-call cycle has five states: proposal, validation, execution, result return, and continuation. We will assume one tool per turn for clarity. Parallel tool calls are a later complication — they multiply the state, not the model.

The Five States of One Tool-Call Cycle

A five-state loop runs from Proposal to Validation to Execution to Result return to Continuation, then back to Proposal. The Validation-to-Execution arrow is highlighted as the side-effect boundary; the other transitions are shown as ordinary flow.
Trace the cycle—and notice that execution is where intent can become an external change.

Each state has inputs, outputs, and transitions out. Name them precisely, because debugging gets much cheaper once you can say which state you are in.

Proposal. The model emits a tool name plus arguments. Nothing has happened yet. This state is pure intent, and it is always safe to discard.

Validation. Your code checks the call against schema, permissions, and preconditions. This is the first place a call can be rejected before any side effect occurs.

Execution. The tool actually runs. This is the only state that can change the outside world.

Result return. The output — or the error — is serialized back into the conversation as a message tied to the call's identifier.

Continuation. The model reads the result and either produces a final answer or proposes another call, closing the loop.

Picture a five-node cycle. Four edges are pure: proposal → validation, validation → execution, execution → result return, result return → continuation. One edge crosses into the world: validation → execution. That single edge is where every dangerous failure lives.

A side effect is any change that survives a retry — a write, a send, a charge, a delete. If running the tool twice leaves the world different from running it once, the tool has a side effect.

Knowledge check

Check your understanding

Answer this question before you continue.

A team wants to mark the point in its tool-call cycle where the outside world can first change. Which state should it mark?
Single Choice

Focus: Identify the state in which a tool can change the outside world.

A Worked Trace: One Call, Five Transitions

Abstract states become observable when you trace one sequence. Take a small tool with a visible side effect: a record lookup followed by an update.

The application state is a messages list, a pending call identifier, and application-side bookkeeping — attempt counts, idempotency keys. Watch it move.

  1. Proposal. The model emits update_record with {id: "42", status: "shipped"}. State: one assistant message containing the call. No external change.
  2. Validation. Your code checks that id is a string, status is an allowed value, and the caller has write permission. All pass. State: unchanged, but the call is now cleared to execute.
  3. Execution. The tool runs. The record in the database now says shipped. State: the world changed. This is the edge that matters.
  4. Result return. The tool returns {ok: true}. Your code appends a tool message tied to the call identifier. State: the conversation now contains the outcome.
  5. Continuation. The model reads {ok: true} and produces a final answer: "Record 42 is marked shipped." The loop reaches a terminal state.

Now rewind to step 2. Suppose validation had rejected the call — status was "shiped", a typo. The conversation would contain a tool message describing the rejection, and the model would propose a corrected call. The loop would have exited early, cheaply, before the world changed. That is the whole point of putting validation before execution: it is the last free exit.

Knowledge check

Check your understanding

Answer this question before you continue.

A proposed update contains the misspelled status `shiped`. Validation rejects it before the tool runs. What should the application infer about the record and the next step?
Scenario Interpretation

Focus: Trace the state consequences of a call rejected during validation.

Where Failures Enter the Machine

Stop treating all errors as one category. Classify them by the state they occur in.

FailureStateWorld changed?Cost
Malformed proposalValidationNoCheap
Validation rejectionValidationNoCheap
Execution errorExecutionMaybeDangerous
Result-handling errorResult returnYesConfusing
Continuation errorContinuationYesLooping

Malformed proposal. The model emits a tool name that does not exist, missing required arguments, or arguments of the wrong type. This fails at validation, before execution. Cheapest possible failure.

Validation rejection. The call is well-formed but not permitted, or a precondition is unmet. Also pre-execution, also cheap.

Execution error. The tool ran and threw, timed out, or returned an error payload. The world may or may not have changed. This is the dangerous ambiguity, and it is the one beginners underestimate.

Result-handling error. The tool succeeded but the result was too large, unparseable, or attached to the wrong call identifier. The world changed; the conversation does not know it.

Continuation error. The model misreads a valid result and loops, or stops when it should have continued.

State the distinction plainly: pre-execution failures are recoverable by construction. Post-execution failures require you to reason about what already happened.

Knowledge check

Check your understanding

Answer this question before you continue.

A tool throws an error during execution. Which conclusion is supported by the state-machine model?
Misconception Check

Focus: Distinguish the uncertainty of an execution error from a pre-execution failure.

Retry Safety: The Only Question That Matters

Idempotent means repeating the operation leaves the system in the same state as doing it once. That one word decides your retry policy.

Retrying a proposal or a validation rejection is safe — no side effect occurred. Retrying execution is safe only if the tool is idempotent or the operation carries an idempotency key.

And here are the assumptions that rule depends on. If any is unstated, the retry is a guess:

  • The tool reports failure honestly.
  • The failure happened before the side effect.
  • No partial write occurred.

Common mistake: retrying on timeout. A timeout tells you the response was lost, not that the work was not done. The server may have completed the charge and simply failed to reply.

The practical pattern is to attach a stable identifier to each call, so a retry can be recognized as the same intent rather than a new one. Before you write retry logic, ask one question: can this tool run twice with the same arguments and produce the same world?

Knowledge check

Check your understanding

Answer this question before you continue.

A write operation timed out, and the application does not know whether the server completed it. Which retry policy best follows the article's guidance?
Comparison Reasoning

Focus: Choose a retry policy that accounts for idempotency and uncertainty about partial execution.

Side Effects and the Boundaries They Create

Side effects split the state machine into two regions. Everything before execution is reversible. Everything after execution may not be.

The practical consequence is a design rule: put validation, permission checks, and argument coercion before execution, because that is the last cheap exit. For irreversible tools, plan a confirmation or dry-run state rather than relying on retries to fix mistakes.

There is a second boundary: the loop itself. A continuation that keeps proposing calls needs an explicit stopping condition, or the machine never reaches a terminal state. Bounding loops is a topic of its own, but the boundary starts here.

When not to use this model: for a single, read-only, always-succeeding tool, the full state machine is overhead. A straight call is fine. The model earns its keep when failure, side effects, or repetition appear.

Reading the Machine in Your Logs

The state machine is also a debugging instrument. Log the state name, the call identifier, the arguments, the outcome, and the attempt count for each transition. With that record, a failure becomes a lookup: which state, which edge, which assumption broke.

Now the opening symptoms resolve. The duplicate charge lives on the execution edge, retried without an idempotency key. The stale result lives in result return, attached to the wrong call identifier. The infinite loop lives in continuation, missing a stopping condition. None of them are mysteries. All of them are locations.

The Decision Rule

Before you ship any tool-calling loop, do three things. Name the five states. Mark the one edge that crosses into the world. For each tool, decide whether it is safe to run twice — and write down the assumptions that answer depends on.

Your next step is concrete: sketch the five-state diagram for one tool in your own project, and label which edges are retry-safe. If you cannot label an edge, that is the edge that will charge the card twice. The natural next topic is bounding loops and stopping conditions in multi-step agents — because once you can see one cycle clearly, the question becomes how many of them you are willing to run.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

An application is adding an irreversible delete tool. Which design best reflects the article's side-effect boundary?
Question 1 of 2Scenario Interpretation

Focus: Apply the side-effect boundary to the design of an irreversible tool workflow.

Logs show that a tool completed successfully, but its result was attached to a different call identifier. The model then acts on stale information. Which state best locates the failure?
Question 2 of 2Scenario Interpretation

Focus: Use logged state and call identifiers to localize a tool-calling failure.

References

  1. Function calling | OpenAI APIdevelopers.openai.com
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.

Close-up of a car dashboard at night with illuminated speedometer and tech displays.
intermediate
10 min read

Build a Simple LLM Agent

Most beginners expect an agent to be a special kind of model—something with built-in magic that can browse the web, run code, and get things done. Then…

Read tutorial