Skip to content
beginner

Practice Diagnosing LLM Output Problems from Generation Settings

The output looks wrong, and your hand is already on the temperature slider. Before you drag it, notice what you are actually doing: guessing. The same…

Published 2026-10-03Updated 2026-10-0411 min read
Close-up of a yellow Ethernet cable with connectors on a blue background.
Close-up of a yellow Ethernet cable with connectors on a blue background. Photo by Ann H on Pexels.

The output looks wrong, and your hand is already on the temperature slider. Before you drag it, notice what you are actually doing: guessing. The same broken-looking answer can come from four unrelated causes, and only one of them lives in the settings panel. This is a drill, not a lecture. You will read request configurations and symptoms, commit to a diagnosis, and let a checker tell you where your reasoning broke.

Why the Settings Slider Is the Wrong First Move

Changing a setting feels productive because it is fast, reversible, and visible. You move a number, you re-run, something different comes back. Motion feels like progress.

But a setting change can only fix a symptom that the setting actually produces. Temperature, top-p, and token limits shape how the model selects tokens and when it stops. They do not shape what the model knows, what your prompt asked for, or whether the evidence you supplied was usable. Reach for a slider when the problem is mechanical. Reach for the prompt when the problem is epistemic.

The trouble is that the symptom rarely tells you which one you have. A cut-off answer, a vague answer, and a confidently wrong answer all look like "the model failed." They are not the same failure.

You will classify every case into one of five buckets:

BucketWhat it meansWhere the fix lives
TruncationOutput stopped before it finishedToken limit
Sampling variationOutput drifts between runsTemperature, top-p
Missing contextThe needed facts were never in the requestPrompt or retrieval
Task/evidence problemFacts were present but the task or evidence was unusablePrompt
Insufficient evidenceThe case does not contain enough to chooseNothing yet — gather more

If you already know what temperature, top-p, and token limits do mechanically, this article is about using that knowledge as a diagnostic instrument. If you need a refresher, the generation-settings reference on this site covers the mechanics; here we only care about what each setting can and cannot repair.

Set Up the Practice Harness

The exercise runs on the Python 3 standard library alone. No pip install, no API key, no network calls. If you can run python3, you can run this.

You need three files in the same folder:

  • diagnose_cases.py — the checker script.
  • cases.json — the supplied request configurations plus the observed output symptoms.
  • answers.json — where you write your diagnoses and rationales.

If your course or site bundle already includes these files, download them into one directory and skip to the run command. If you do not have them yet, create them yourself. The script is short, and building it is part of the lesson: you cannot diagnose a system you have not looked inside.

Here is a minimal diagnose_cases.py you can save as-is:

import json
import sys

def load(path):
    with open(path, "r", encoding="utf-8") as f:
        return json.load(f)

def main():
    cases = {c["id"]: c for c in load(sys.argv[1])}
    answers = {a["id"]: a for a in load(sys.argv[2])}
    hits, misses = 0, 0
    for cid, case in cases.items():
        answer = answers.get(cid)
        if not answer:
            print(f"{cid}: no answer submitted")
            misses += 1
            continue
        expected = case["expected"]
        ok = answer["diagnosis"] == expected["diagnosis"]
        if ok:
            hits += 1
            print(f"{cid}: correct ({answer['diagnosis']})")
        else:
            misses += 1
            print(f"{cid}: expected {expected['diagnosis']}, got {answer['diagnosis']}")
            print(f"    feedback: {expected['feedback']}")
    print(f"\n{hits} correct, {misses} to review")

if __name__ == "__main__":
    main()

And here is a tiny cases.json with two cases so you can see the shape before you write your own:

[
  {
    "id": "case-01",
    "config": {"temperature": 0.9, "top_p": 1.0, "max_tokens": 200},
    "task": "Summarize the attached article in three paragraphs.",
    "symptom": "Output ends mid-sentence in paragraph two.",
    "expected": {
      "diagnosis": "truncation",
      "feedback": "Output is cut, not concluded. The token limit is low relative to the requested length."
    }
  },
  {
    "id": "case-02",
    "config": {"temperature": 0.7, "top_p": 1.0, "max_tokens": 800},
    "task": "Answer the question using only the supplied document.",
    "symptom": "Confident, specific answer that cites facts not present in the document.",
    "expected": {
      "diagnosis": "missing context",
      "feedback": "The needed facts were never in the request. No setting supplies them."
    }
  }
]

Your answers.json mirrors that structure with your own judgment:

[
  {"id": "case-01", "diagnosis": "truncation", "rationale": "..."},
  {"id": "case-02", "diagnosis": "missing context", "rationale": "..."}
]

Run the checker:

python3 diagnose_cases.py cases.json answers.json

The checker compares your diagnosis against expected feedback. It grades your reasoning, not your vocabulary. A right label with a wrong rationale is still a miss, because the label is the cheap part and the reasoning is the skill.

What success looks like on the first pass is not a perfect score. It is a visible list of which cases you misread and why. That list is the actual output of this exercise.

Note: The checker reads the expected field from cases.json. If you are writing your own cases, fill that field honestly — a case with a guessed expected answer teaches you nothing.

Read the Case Before You Read the Settings

A decision path starts with whether the answer is cut off, then checks whether repeated runs differ, and finally whether the needed facts are present. The outcomes are truncation, sampling variation, missing context, task or evidence problem, or insufficient evidence.
Follow the symptom and evidence before choosing an intervention; only truncation and sampling variation point first to generation settings.

Each case record contains three things: the request configuration, the prompt or task description, and the observed output symptom.

Read them in that order — symptom first, config second, hypothesis third. If you read the config first, you start hunting for a setting to blame instead of asking what the output is telling you. That is how a missing-document problem gets "fixed" by lowering temperature, which changes nothing except your confidence.

Ask three questions of every symptom:

  1. Is content missing at the end? If the output stops mid-sentence or mid-structure, truncation is a candidate.
  2. Does the answer change between runs? If the same request produces different wording, ordering, or detail selection, sampling variation is a candidate.
  3. Is the needed information present anywhere in the request? If it is absent, no setting will conjure it. If it is present but the output ignored or misused it, the problem is in the task or the evidence.

Here is a worked read-through. Suppose a case shows a request asking for a three-paragraph summary of a supplied article, with a token limit set, and the observed output is two and a half paragraphs that end mid-word. Symptom first: content is missing at the end, and it is cut rather than concluded. Config second: the token limit is low relative to the requested length. Hypothesis: truncation. The confirming test is to run it again — truncation repeats, because the limit is a hard stop, not a dice roll.

Knowledge check

Check your understanding

Answer this question before you continue.

A request asks for a three-paragraph summary. The output ends mid-sentence, the token limit is low for that length, and another run ends at a similar cutoff. Which diagnosis best fits?
Scenario Interpretation

Focus: Classify an output that ends abruptly by using the symptom and request configuration as evidence.

Truncation vs. Sampling Variation

These two are the setting-adjacent buckets, and beginners confuse them because both feel like "the model stopped short of what I wanted."

Truncation signals: the output ends mid-sentence or mid-structure, the ending looks cut rather than concluded, and the symptom is stable across repeated runs.

Sampling variation signals: the same request produces different answers on different runs, and the differences are in wording, ordering, or detail selection rather than in missing content.

The discriminating test is simple: run it again. Truncation repeats. Sampling variation drifts.

The bounded interventions follow from the mechanism. For truncation, raise the token limit — you are giving the model more room to finish. For variation, lower temperature or top-p — you are narrowing the pool of tokens the model samples from.

Now the part that matters more than the fix: what each one cannot do. A higher token limit will not make a wrong answer right; it will just let the model be wrong for longer. A lower temperature will not add information the model never had; it will make the model more consistent about not having it.

Common mistake: Treating low temperature as a correctness dial. It is a consistency dial. Consistent wrong answers are still wrong.

Knowledge check

Check your understanding

Answer this question before you continue.

The same request produces complete answers that vary in wording and detail selection across runs. Which intervention best matches this diagnosis, and what limitation should you keep in mind?
Comparison Reasoning

Focus: Choose a bounded setting intervention for sampling variation and recognize its limit.

Missing Context vs. Task and Evidence Problems

These are the non-setting buckets, and the boundary between them is where this exercise's real lesson lives.

Missing context: the request never supplied the facts the answer needed — no document, no data, no prior turn. The model filled the gap, which is why the output reads confidently. Fluent wrongness is the signature of a model doing its best with nothing.

Task/evidence problem: the information was present, but the instruction was ambiguous, the format was underspecified, or the supplied evidence contradicted itself. The model had what it needed and still could not use it cleanly.

The distinguishing question: was the needed information absent from the request, or present but unusable?

Neither one yields to a setting change. Temperature, top-p, and token limits shape how the model selects and stops. They do not shape what the model knows. If the evidence was never in the request, no slider will put it there.

Then there is the honest bucket: insufficient evidence. Sometimes a case does not contain enough information to choose a diagnosis. Saying so is the correct answer, not a cop-out. A diagnosis you cannot support is a guess wearing a label.

Tip: If you find yourself arguing for a bucket with no specific evidence from the case, you are probably in the insufficient-evidence bucket and resisting it.

Knowledge check

Check your understanding

Answer this question before you continue.

A request includes the relevant figures, but two supplied figures contradict each other and the instruction does not say which source to trust. Which bucket best describes the problem?
Scenario Interpretation

Focus: Distinguish an evidence or task problem from missing context.

Write a Rationale the Checker Can Grade

A correct label is half the work. The rationale is the other half, and it has three parts:

  1. Name the bucket.
  2. Cite the specific evidence in the case that supports it.
  3. State what your chosen intervention cannot fix.

The third part matters most. It is the difference between "I changed a setting" and "I understood what the setting controls." Anyone can move a slider. Naming the boundary of what that slider can do is the skill this exercise is built to train.

Here is a weak rationale and a strong one for the same case — a request that supplied no source document and produced a confident, specific, unsupported answer:

Weak: "The answer is wrong, so I lowered the temperature to make it more accurate."

Strong: "This is missing context. The request contains no source document, and the output asserts specifics that appear nowhere in the prompt. Lowering temperature cannot fix this because the model never had the facts; it would only make the same unsupported answer more consistent."

The weak version restates the symptom and proposes an intervention that does not match the diagnosis. The strong version names the bucket, points at the evidence, and closes the door on the wrong fix.

Knowledge check

Check your understanding

Answer this question before you continue.

A case has no source document, yet the output makes specific unsupported claims. Which rationale includes the three elements the article says the checker should be able to grade?
Single Choice

Focus: Construct a rationale that supports a diagnosis and states the intervention's limitation.

Edit a Case and Diagnose Again

Now the experiment. Pick one case, change a single field in cases.json, and predict what the diagnosis should become before you re-run the checker.

Two edits are worth trying:

Raise the token limit on a truncation case. The symptom should resolve — the output finishes. You have confirmed that the limit was the cause, not a coincidence.

Remove the source document from a missing-context case. The symptom stays. The model still produces fluent output; it just produces the wrong fluent output. No setting rescues it, because the evidence is gone. This is the edit that proves the core lesson — settings cannot supply missing evidence.

After each edit, read the checker's feedback carefully. Did the diagnosis change for the reason you predicted, or did you change the case in a way that muddied the signal?

Warning: Change one variable at a time. Change two fields and you learn nothing about either.

When to Reach for a Setting, and When Not To

Compress the whole exercise into a rule you can carry into real work.

Reach for a setting when the symptom is mechanical: output cut off, output unstable across runs, output repetitive.

Do not reach for a setting when the symptom is epistemic: the answer is wrong, unsupported, or built on information that was never supplied.

The one-line version: settings control how the model chooses and stops; they do not control what the model knows.

That rule also tells you where diagnosis sits in your workflow. Diagnosis comes before prompt rewriting, and both come before touching a slider. If you skip the diagnosis, you are optimizing a system you have not understood.

One limit worth stating plainly: this exercise uses supplied cases with expected feedback. It teaches classification discipline, not a guarantee about any specific production model's behavior. Real systems are messier, and the buckets are a starting frame, not a verdict.

Before you touch a setting, name the symptom, name the bucket, and ask whether the missing thing is a mechanism or a fact. If it is a mechanism, the slider is yours. If it is a fact, the slider was never going to help.

Your next move: re-run the exercise with your own real request configurations and symptoms. Take three outputs that disappointed you this week, write them up as cases, and diagnose them cold. The debugging workflow on this site picks up where the non-setting cases lead — task ambiguity, missing context, format failure, and unsupported claims — and turns the classification habit into a repeatable repair process.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

You increase only the token limit in a case where the output repeatedly ends mid-sentence. If the limit was causing the cutoff, what result should you predict?
Question 1 of 2Output Prediction

Focus: Predict the diagnostic effect of changing one relevant case field at a time.

A model gives a fluent, unsupported answer because the request omitted the source document. Which response best follows the article's overall diagnostic rule?
Question 2 of 2Misconception Check

Focus: Apply the distinction between mechanical generation settings and missing evidence.

References

  1. Optimizing LLM Accuracy | OpenAI APIdevelopers.openai.com
  2. Understanding Temperature, Top P, and Maximum Length in LLMslearnprompting.org
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.