Skip to content
beginner

Practice Designing a Structured-Output Contract

The model returns clean JSON. Every brace closes. Every key is present. And the answer is still wrong.

Published 2026-10-03Updated 2026-10-049 min read
Close-up of intricate sand patterns captured in monochrome, highlighting texture and form.
Close-up of intricate sand patterns captured in monochrome, highlighting texture and form. Photo by Warren Adler on Pexels.

The model returns clean JSON. Every brace closes. Every key is present. And the answer is still wrong.

That is the moment most beginners learn the hard truth: valid shape is not the same thing as a valid answer. A response can pass every format check and still be useless, because the field was filled with a plausible guess, the category was technically allowed but factually wrong, or a required field got stuffed with filler just to satisfy the shape.

If you have already learned how to ask for structured output and why validation still matters, this is the next rep. Here you will design a structured-output contract for one small task, then test it against four messy inputs before you ever judge the model. The goal is to decide what a response must mean, not just how it is shaped.

Why Valid JSON Is Not a Contract

A two-by-two matrix compares format validity with semantic correctness. Only output that passes both checks is accepted; valid format with incorrect meaning still fails.
Check both the response’s shape and what its values mean; passing one check does not guarantee passing the other.

A schema describes shape. A contract describes meaning. Those are two different layers, and beginners collapse them constantly.

Format requirements cover the mechanical half: which keys exist, what types they hold, which are required, whether the JSON parses. Semantic requirements cover the judgment half: what values are allowed to mean, when a value is a guess instead of a fact, and what the response should say when the input does not fit.

The contract is the written artifact that covers both. The schema is only the format half of it.

Here is the failure in miniature. Imagine a task that classifies a support message. The model returns:

{
  "category": "billing",
  "priority": "high",
  "reason": "Customer seems upset about a charge."
}

Every field is present, every type is correct, the JSON is valid. But if the message was actually about a login failure, the category is a confident lie wrapped in perfect syntax. A validator would pass it. A human would not.

Common mistake: Treating "the JSON parsed" as "the answer is correct." Those are two separate checks, and you need both.

Knowledge check

Check your understanding

Answer this question before you continue.

A support-message response contains every required key, uses the required types, and parses as JSON, but labels a login failure as billing. What can you conclude from the format checks alone?
Misconception Check

Focus: Distinguish format validity from semantic correctness when evaluating a structured response.

Pick a Bounded Task and Name Its Decisions

Before you write a single field, pick a task with a narrow input space. Narrow means the input is short, the possible outputs are limited, and you can imagine the whole decision space in your head.

A good starter: classify a short support message into a category, a priority, and a one-line reason.

Now do the part beginners skip. List the decisions a downstream reader or program must make from the output. Each decision implies a field.

  • "Which team should see this?" → needs a category.
  • "Does it jump the queue?" → needs a priority.
  • "Why did the model choose that?" → needs a short reason.

That is three fields. Resist adding more. A field that drives no decision, no join, and no stored artifact is an invitation for creative filler. If nothing downstream reads it, cut it.

Finally, state the task boundary in one sentence: what is in scope and what is out. For this task: in scope is a short customer message about a product or service issue; out of scope is anything that is not a support message at all — a sales question, spam, or a message in a language the system does not handle.

Knowledge check

Check your understanding

Answer this question before you continue.

A team classifies short support messages and only needs to route each message, determine urgency, and see a brief justification. Which design choice best follows the article's field-selection method?
Scenario Interpretation

Focus: Choose fields by identifying which downstream decisions the output must support.

Write the Contract: Fields, Types, and Allowed Values

Draft the contract in plain language first. Then translate it into a schema-like shape. Plain language forces you to say what each field means; the schema shape forces you to say what it is.

FieldTypeRequiredMeaning
categorystring, closed setyesThe single best-fit issue area
prioritystring, closed setyesHow urgently a human should look
reasonstring, shortyesOne sentence of evidence from the message
statusstring, closed setyesWhether the message could be classified at all

Now define the allowed values as closed sets wherever the decision space is closed. For category: billing, technical, account, other. For priority: low, medium, high. For status: classified, ambiguous, incomplete, out_of_scope.

Why closed sets? An open string invites drift. Ask for a category as free text and you will get billing, Billing Issue, billing-related, and payment problem — four labels for one idea, and now your downstream code has to guess which ones match.

Keep required fields minimal. A field that is sometimes unknown should either be optional or carry an explicit unknown value rather than being forced into a guess.

Order matters too. Put evidence before conclusion. If reason comes before category, the model writes its reasoning first and then commits to a label that fits. If the label comes first, the reasoning tends to justify whatever it already picked.

Note: Many structured-output implementations require every field to be declared required. When that is your constraint, unknown becomes an allowed value rather than an omitted key. Design for the tool you actually have.

Knowledge check

Check your understanding

Answer this question before you continue.

The contract allows `billing`, `technical`, `account`, and `other` for `category`. The model returns `payment problem` for a duplicate-charge message. Which diagnosis best fits this output?
Debugging

Focus: Identify an allowed-value violation separately from a semantically similar value.

Decide What Happens to Messy Inputs

This is the section most beginner contracts omit, and it is where real reliability lives. You need three explicit dispositions: ambiguous, incomplete, and out-of-scope.

For each one, decide the contract's answer before you see any model output. Every disposition must still return all four required fields, so the response stays complete and testable:

  • Ambiguous — the message could reasonably fit two categories. Set status to ambiguous, pick the most likely category, set priority to your best guess, and note the competing reading in reason.
  • Incomplete — the message is too short to classify. Set status to incomplete, set category to other, set priority to low, and use reason to say what detail is missing.
  • Out-of-scope — the message is not a support issue. Set status to out_of_scope, set category to other, set priority to low, and use reason to name why it falls outside the task.

The key move is that ambiguity handling lives in the contract, not in a hopeful instruction like "ask for clarification if unsure." A vague instruction gives the model room to silently pick a side. A named status value forces it to admit the input was messy.

Out-of-scope inputs especially need a named exit. Without one, the model is forced to cram a sales question into billing or technical, because those are the only doors you left open.

Knowledge check

Check your understanding

Answer this question before you continue.

Using the article's stated contract, what should the response do with the input “It's broken.”?
Output Prediction

Focus: Apply the specified disposition for an input too incomplete to classify.

Test the Contract Against Four Examples

Now the deliberate part. Here is a small test set. Hand-fill the contract for each input yourself before you run anything. Write down what you believe the correct output is, field by field. If you cannot fill it consistently, the contract is underspecified — not the model.

Input 1 (clean): "I was charged twice for my subscription this month. Please refund the duplicate."

Input 2 (ambiguous): "I can't log in and my invoice looks wrong."

Input 3 (incomplete): "It's broken."

Input 4 (out-of-scope): "Do you offer enterprise pricing for a team of 200?"

Here is the intended contract output for each:

[
  {
    "category": "billing",
    "priority": "high",
    "reason": "Customer reports a duplicate charge and requests a refund.",
    "status": "classified"
  },
  {
    "category": "account",
    "priority": "medium",
    "reason": "Message could fit account (login) or billing (invoice); login chosen as primary.",
    "status": "ambiguous"
  },
  {
    "category": "other",
    "priority": "low",
    "reason": "No issue detail provided; cannot classify.",
    "status": "incomplete"
  },
  {
    "category": "other",
    "priority": "low",
    "reason": "Sales inquiry, not a support issue.",
    "status": "out_of_scope"
  }
]

Now run the model and compare its output against your hand-filled version, field by field. Classify each mismatch into one of four buckets:

  • Format failure — missing key, wrong type, broken JSON.
  • Allowed-value violation — a value outside your closed set.
  • Semantic error — a valid value that is simply wrong.
  • Contract gap — the model did something reasonable that your contract never anticipated.

That last bucket is the valuable one. It means your specification, not the model, needs work. Revise the contract for gaps and leave the model unchanged. One variable at a time.

Common Contract Mistakes and How to Spot Them

Run your draft against this list before you trust it.

  • Too many required fields. The model invents values to satisfy the shape. If a field is often unknowable, it should not be required.
  • Open strings where a closed set was available. Output drifts into near-duplicate categories.
  • No disposition for ambiguity. The model silently picks a side and hides the uncertainty.
  • Incomplete dispositions. A rule that says "stop" without specifying every required field leaves the response untestable.
  • Confusing format compliance with correctness. A passing validator is not a passing answer.
  • Writing the contract after seeing the output. Now it describes what the model did, not what you need. Write it first.

When a Contract Is Worth Writing — and When It Is Not

Write a contract when the output feeds code, a database, a downstream decision, or another person who needs consistent fields. Write one when you keep cleaning up the same mess by hand, when categories come back inconsistent, or when two people argue about what a field "should" contain.

Skip the ceremony for one-off exploratory prompts where you are still discovering what you even want to ask. A contract locks in decisions. Locking in decisions before you have made them just freezes your confusion.

The payoff is that a contract is a reusable asset. Once written, it survives model changes, prompt rewrites, and new examples. The prompt is disposable; the contract is the part that compounds.

Your Next Rep

Take the contract you just wrote and add one new allowed value to a closed set — say, a fourth category. Re-run the same four examples and watch which outputs shift. You will usually learn more from that one change than from a dozen new prompts.

Or swap the task entirely and write a second contract from scratch in under ten minutes using the same four steps: name the decisions, define the fields and allowed values, decide the messy-input dispositions, then hand-fill four examples before running the model.

Keep the contract file next to the prompt. The specification and the request should travel together, so the next person — including future you — can see what the response was always supposed to mean.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

While testing a new representative input, the model returns a response that seems reasonable, but the contract never says how that kind of input should be handled. What is the best next step?
Question 1 of 2Scenario Interpretation

Focus: Distinguish a contract gap from an output error when testing representative inputs.

A team is still exploring what information it wants from a one-off prompt and has not settled on the decisions its output should support. Based on the article, which approach is most appropriate now?
Question 2 of 2Comparison Reasoning

Focus: Decide when a structured-output contract is useful based on whether output decisions are settled and reused.

References

  1. Structured model outputs | OpenAI APIdevelopers.openai.com
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.

A 3D rendering of a neural network with abstract neuron connections in soft colors.
beginner
8 min read

Common Prompting Mistakes

You write a prompt. You press enter. The model replies with something generic, slightly off, or completely wrong. So you rewrite, try again, and get…

Read tutorial