Skip to content
intermediate

Practice Checking LLM Tool Results Against the Application Contract

A tool call returns status: "ok". The pipeline moves on. Nobody notices that the result answers a slightly different question than the one the model asked.

Published 2026-10-03Updated 2026-10-0410 min read
Detailed texture of sand on a beach in Antalya, capturing natural patterns and earthy tones.
Detailed texture of sand on a beach in Antalya, capturing natural patterns and earthy tones. Photo by Andrew Schwark on Pexels.

A tool call returns status: "ok". The pipeline moves on. Nobody notices that the result answers a slightly different question than the one the model asked.

That is the failure this exercise trains you to catch. Not crashes. Not timeouts. The quiet, well-formed, plausible result that is wrong in a way only the contract can reveal.

You already know how a tool call travels from model to application and back. This drill picks up at the step most tutorials skip: after execution, before trust. You will run a small standard-library checker against static records, decide a disposition for each case, and defend it from the record itself. Then you will change one thing and decide again.

Why a Successful Call Is Not a Valid Result

Three different questions hide inside the phrase "the tool call worked."

  1. Did the call execute without an error?
  2. Did it return the operation that was actually requested?
  3. Does the result satisfy the contract the application requires?

Most pipelines answer question one and assume the rest. That assumption is the weak model. It survives demos because demos use clean inputs and obvious tools. It fails in production the moment an argument gets swapped, a unit gets misread, or a partial result gets reported as complete.

The classic shape is argument inversion. A currency tool is asked for the fluctuation of EUR against USD. The call is well-formed JSON. The tool executes. The result is internally consistent. And it is wrong, because base and symbols were reversed, so the tool computed USD against EUR instead. Nothing in the execution layer can catch that. Only a contract that says "the base argument must match the subject of the user's request" can.

So we replace the weak model with a stronger one: a tool result is a claim, and the contract is what turns a claim into a decision.

That decision has exactly four allowed outputs:

DispositionMeaning
AcceptResult matches the requested operation and every contract clause.
RejectResult is wrong, malformed, or answers a different operation.
Bounded retryThe contract names this failure class as retryable and caps attempts.
EscalateThe contract is silent, ambiguous, or the failure is outside the tool's authority.

Four outputs, no vibes. Every case forces a decision.

One rule governs the whole exercise: retry is never a default. It is available only when the supplied contract authorizes it. If the contract is silent, the safe move is reject or escalate.

Knowledge check

Check your understanding

Answer this question before you continue.

A currency call executes successfully, but its `base` and `symbols` are reversed from the requested operation. What disposition fits the article's rule?
Scenario Interpretation

Focus: Distinguish successful execution from a result that matches the requested operation.

The Three Record Types and What Each One Proves

Three inputs—proposed call, execution result, and contract—converge on a validation step, which leads to accept, reject, bounded retry, or escalate.
A successful execution is only one input: compare the requested operation and returned result against the contract before choosing a disposition.

Before you run anything, understand the data model, or the checker's output will read like noise.

Proposed call record. The tool name, the arguments, and the user intent the call claims to serve. This is what the model asked for.

Execution result record. Status, payload, error field, timing, and any partial-completion signal. This is what came back.

Contract record. Required fields, allowed value ranges or enums, argument-to-result correspondence rules, retry policy, and escalation conditions. This is what the application demands.

Picture three columns: call, result, contract. The disposition decision sits at the intersection of all three. A result can be valid against the contract and still answer the wrong operation. A call can be well-formed and still violate the contract. You need all three records to decide.

The contract is the only source of authority for retry. That sentence is the spine of the exercise. When the contract is silent on a failure class, you do not get to invent a retry policy on the fly. You reject or you escalate.

Knowledge check

Check your understanding

Answer this question before you continue.

Which set of records does the article say you need to decide whether a result is acceptable?
Comparison Reasoning

Focus: Identify the evidence each of the three record types contributes to a disposition decision.

Run the Checker: Setup and First Output

Dependencies are minimal by design: Python 3 standard library only. The script uses json, argparse, and sys. No network, no API keys, no external services. The records are static files, so the exercise is reproducible and the reasoning is the only variable.

Run it:

python3 check_contracts.py records.json contracts.json answers.json

The inputs:

  • records.json holds paired proposed-call and execution-result entries, keyed by case id.
  • contracts.json holds the per-tool contract.
  • answers.json holds your submitted disposition and reason for each case.

The output gives you three things per case: pass or fail on the disposition, a diff between your reason and the expected feedback, and a summary count at the end.

Note: The checker compares reasons, not just labels. A right disposition for the wrong reason will not survive the next case, because the reasoning is what generalizes. The label is just the residue.

Success is correct baseline reasoning on the supplied cases. Not a perfect score on the first run. If you get a case wrong, the diff tells you which field or clause you missed, and that is the actual lesson.

Reading a Case: From Record to Disposition

Walk one case end to end. The pattern repeats for every case after it.

Start with the proposed call and restate the requested operation in plain language. "Get the EUR/USD fluctuation for 2020." Write it down. This is your anchor.

Then check the result against the contract, field by field:

  • Presence. Are all required fields there?
  • Type. Does each field have the expected type?
  • Range. Do numeric values fall inside allowed bounds? Do enums match allowed values?
  • Correspondence. Does the result actually correspond to the arguments? If the call asked for EUR as the base, does the payload reflect EUR as the base?

Now classify what you found.

A contract violation — a missing field, an out-of-range value, a swapped argument — is a reject. Do not silently repair it. Repairing hides the bug and teaches the model that malformed calls work.

A transient execution failure the contract explicitly permits retrying is a bounded retry. The contract must name the failure class and cap the attempts. State the cap when you justify.

A case the contract cannot resolve — ambiguous intent, missing authority, conflicting rules — is an escalate. Hand it off with the record attached.

The reason string that earns a pass names the specific field, rule, or clause that drove the decision. "Looks wrong" fails. "base is usd but the requested operation names EUR as the subject, violating the correspondence rule" passes.

A short decision tree:

Contract satisfied?
  yes -> Accept
  no  -> Contract authorizes retry for this failure class?
           yes -> Bounded retry (state the cap)
           no  -> Contract resolves the ambiguity?
                    yes -> Reject
                    no  -> Escalate

Knowledge check

Check your understanding

Answer this question before you continue.

A case has a transient execution failure. Its contract explicitly permits retrying this failure class, with a maximum of two attempts. Which disposition and justification fit?
Scenario Interpretation

Focus: Apply a contract-authorized retry policy while respecting its attempt cap.

The Four Dispositions and Where Each One Belongs

Accept when the result matches the requested operation and every contract clause. No further action. This is the only disposition that lets the pipeline continue without a note.

Reject when the result is wrong, malformed, or answers a different operation. Do not silently repair it. A rejected result is evidence about the call, the tool, or the contract, and that evidence is worth keeping.

Bounded retry only when the contract names the failure class as retryable and caps the attempts. State the cap and the backoff expectation in your reason. If the contract says "retry up to 3 times on timeout," that is authority. If the contract says nothing, you have none.

Escalate when the contract is silent, ambiguous, or the failure is outside the tool's authority. Hand off with the record attached so the next person has the same evidence you did.

Two mistakes account for most wrong answers.

Common mistake: Treating retry as the polite default. Retrying a logic error burns budget and hides the bug. A swapped argument will be swapped again on the next attempt. Retry is for transient failures the contract anticipated, not for confusion.

Common mistake: Accepting a result because the payload is well-formed JSON. Format validity is not contract validity. A perfectly shaped object can still answer the wrong question.

Failure Modes the Checker Will Catch

When you fail a case, it is usually one of these.

Argument inversion. The call is well-formed but the parameters are swapped. The result is internally consistent and externally wrong. This is the hardest one to catch by eye, because nothing looks broken.

Status laundering. The execution result reports success while the payload shows a partial or empty outcome. The status field says ok; the payload says "3 of 10 records returned." The contract decides which one wins.

Contract drift. The record was written against an older contract version, so the rule you are applying no longer exists. Check the contract version before you apply a clause.

Reason mismatch. Correct disposition, vague justification. The checker flags it because the reasoning will not generalize to the next case.

Over-escalation. Escalating a case the contract clearly resolves. This is a different failure than under-escalating, and it is just as costly: it pushes work onto a human who did not need to be involved.

Change One Thing and Decide Again

The exercise is not memorizing dispositions. It is proving your reasoning transfers. Change one variable, predict the new disposition, then run.

Experiment A: flip one argument. Take a case you accepted. Swap base and symbols in the proposed call. Predict the new disposition before running. If you predicted reject, you understood the correspondence rule. If you predicted accept, you were pattern-matching the old answer.

Experiment B: add a retry clause. Find a case you rejected for a transient failure. Add a retry clause to the contract that names that failure class and caps attempts at two. The case becomes a bounded retry. Watch how the same result flips disposition purely because the contract changed.

Experiment C: tighten a range. Take a clean accept. Narrow the allowed range in the contract so the returned value now falls outside it. The case becomes a reject, because the contract no longer permits the value the tool returned. This is the cleanest way to see that the same result can be valid under one contract and invalid under another.

For each change, write the reason string first, then run, then compare. The prediction is the exercise. The checker is just the referee.

Note: This drill covers static records only. Live tool execution, concurrency, and cost accounting are out of scope. Those matter in production, but they are separate problems from the one you are practicing here: deciding whether a finished result matches the requested operation and the contract.

Knowledge check

Check your understanding

Answer this question before you continue.

You take a previously accepted currency case and swap `base` and `symbols` in the proposed call. Under the article's correspondence rule, what should you predict?
Output Prediction

Focus: Predict how a changed proposed-call argument affects contract correspondence and disposition.

What to Carry Forward

A tool result is a claim. The contract is the only thing that turns a claim into a decision. Everything else — the status field, the well-formed JSON, the plausible-looking payload — is testimony, not proof.

The four dispositions are not a scoring rubric. They are a decision rule you can reuse outside this exercise. Accept, reject, retry, escalate. Retry only with authority. Escalate when the contract cannot resolve the case. Reject when it can.

Your next move: take one real tool call from a system you are building. Write down the contract it should satisfy — required fields, correspondence rules, retry policy, escalation conditions. Then run the same four-way decision on the result before you trust it. If you cannot write the contract, you have found the actual gap.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

An execution record says `status: "ok"`, but its payload reports 3 of 10 records returned. The contract requires all 10 and provides no retry authorization. Which disposition follows?
Question 1 of 2Scenario Interpretation

Focus: Use the contract rather than a success status to classify a partial result.

You submit the correct disposition label, but justify it only with “looks wrong.” What should you expect from the checker?
Question 2 of 2Misconception Check

Focus: Explain why a correct disposition label still needs a contract-specific reason.

References

  1. Closing the Data–Training Loop for Robust LLM Tool Callsaclanthology.org
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.