Skip to content
intermediate

Practice Studying LLM Repeatability Across Repeated Runs

You run the same prompt five times and get five answers that mean roughly the same thing but never match word for word. Now what? You cannot tell whether…

Published 2026-10-03Updated 2026-10-0410 min read
Aerial photo showcasing sand dunes and vibrant greenery in a desert landscape.
Aerial photo showcasing sand dunes and vibrant greenery in a desert landscape. Photo by Quang Nguyen Vinh on Pexels.

You run the same prompt five times and get five answers that mean roughly the same thing but never match word for word. Now what? You cannot tell whether the system is unstable or whether you simply have not looked closely enough. That gap between seeing variation and understanding it is where most reliability judgments go wrong.

Here is the stronger model: repeatability is not a vibe you get from a few outputs. It is a property you measure per case, under a fixed protocol, before you are allowed to conclude anything about the task. This article gives you a small runnable fixture and, more importantly, a way to read what it produces.

If you have not yet separated correctness from dependability, skim the idea that one good answer proves little, then come back. This piece assumes you already accept that outputs vary and that a single success is not evidence.

What a Repeatability Test Actually Measures

Before any code runs, fix the mental model. A repeatability test measures the closeness of agreement of repeated measurements taken under identical conditions: same model, same prompt, same generation parameters, same case. Change any of those conditions and you have moved into a wider claim — reproducibility — where the question becomes whether the system still agrees when prompts, versions, or environments shift.

Two layers of variation matter, and conflating them causes bad decisions:

  • Semantic variation — does the meaning shift across runs?
  • Surface variation — does the wording shift while the meaning holds?

A run can be semantically stable but lexically noisy ("The diagnosis is meningitis" versus "Diagnosis is meningitis"), or the reverse: nearly identical wording that quietly changes a decision. The first is usually harmless. The second is the one that bites you in production.

The unit of analysis is the case, not the run. You measure per-case stability, then aggregate across cases. And state the boundary up front: a handful of runs on a handful of cases can reveal instability, but it can never certify stability. Absence of observed variation is not evidence of determinism.

Common mistake: Treating "all five outputs were identical" as proof the system is deterministic. Caching, a temperature of zero, or an under-specified case can all produce identical outputs while hiding real variation.

Knowledge check

Check your understanding

Answer this question before you continue.

Which setup measures repeatability as described in the article?
Single Choice

Focus: Distinguish repeatability testing under fixed conditions from broader reproducibility claims.

The Fixture: Cases, Protocol, and Fixed Parameters

The smallest useful setup is deliberately boring. Use 4–6 short cases with a checkable expected property, and 5 runs per case. A case with a checkable property — a required field, a required format, a required fact drawn from supplied context — is worth far more than a case you can only "feel" about.

Here is a concrete fixture you can run without any API credentials. It uses a tiny local stub that returns one of several plausible outputs per case, so you can practice the full protocol — collection, per-case summary, and interpretation — before you point it at a real model.

import json, random

# A disclosed fixture: 4 cases, each with a checkable expected property.
CASES = [
    {
        "id": "extract_email",
        "text": "Pull the email address from: 'Contact Ana at ana@example.com for details.'",
        "expected": "ana@example.com",
    },
    {
        "id": "classify_sentiment",
        "text": "Classify sentiment as positive or negative: 'The update broke my workflow.'",
        "expected": "negative",
    },
    {
        "id": "format_json",
        "text": "Return JSON with keys name and age for: 'Mira, 34'.",
        "expected": '{"name": "Mira", "age": 34}',
    },
    {
        "id": "extract_date",
        "text": "Extract the date from: 'The invoice is due on March 3, 2026.'",
        "expected": "March 3, 2026",
    },
]

# A stub client that simulates run-to-run variability.
# Each case has a small pool of plausible outputs; some are correct, some are not.
STUB_OUTPUTS = {
    "extract_email": ["ana@example.com", "ana@example.com", "ana@example.com", "ana@example.com", "ana@example.com"],
    "classify_sentiment": ["negative", "negative", "negative", "mixed", "negative"],
    "format_json": ['{"name": "Mira", "age": 34}', '{"name": "Mira", "age": 34}', '{"name": "Mira", "age": 34}', '{"name": "Mira", "age": 34}', '{"name": "Mira", "age": 34}'],
    "extract_date": ["March 3, 2026", "03/03/2026", "March 3, 2026", "2026-03-03", "March 3, 2026"],
}

class StubClient:
    def generate(self, model, prompt, **params):
        # Deterministic per (case, run) so the fixture is reproducible.
        case_id = next(c["id"] for c in CASES if c["text"] == prompt)
        run_index = params.get("_run_index", 0)
        text = STUB_OUTPUTS[case_id][run_index % len(STUB_OUTPUTS[case_id])]
        return type("R", (), {"text": text})()

def run_fixture(cases, client, model, params, runs_per_case=5):
    rows = []
    for case in cases:
        for run_index in range(runs_per_case):
            try:
                response = client.generate(
                    model=model,
                    prompt=case["text"],
                    _run_index=run_index,
                    **params,
                )
                rows.append({
                    "case_id": case["id"],
                    "run_index": run_index,
                    "output": response.text,
                    "error": None,
                })
            except Exception as exc:
                rows.append({
                    "case_id": case["id"],
                    "run_index": run_index,
                    "output": None,
                    "error": repr(exc),
                })
    return rows

MODEL = "stub-v1"
PARAMS = {"temperature": 0.7}

rows = run_fixture(CASES, StubClient(), MODEL, PARAMS)
with open("results.jsonl", "w") as f:
    for row in rows:
        f.write(json.dumps(row) + "\n")

Expected observable output: a results.jsonl file with 20 rows — one per (case, run) pair. If you cannot point to a row for every case-run combination, your run is incomplete.

To swap in a real model, replace StubClient with a client that exposes a generate(model, prompt, **params) method returning an object with a .text attribute. Pin the model identifier and parameters in a run manifest alongside the results.

Debugging signals worth reading literally:

  • Identical outputs across all runs may mean caching or a temperature of zero masking real variation.
  • Wildly different output lengths often mean the case is under-specified.
  • Missing rows mean the error path swallowed a failure.

Knowledge check

Check your understanding

Answer this question before you continue.

The disclosed fixture has four cases and runs each case five times. How many result rows should a complete run produce?
Output Prediction

Focus: Determine whether a repeated-run fixture collected a complete set of case-run observations.

Reading the Results: Within-Case Variation vs Across-Case Failure

A matrix with four case rows and five run columns. Email and JSON rows show five pass marks; sentiment shows four passes and one failure; date shows three passes and two format failures. Row patterns reveal per-case stability, while a marked run column highlights a single-run deviation.
Read across each case to find variation; compare cases only after their individual patterns are clear.

This is the analytical move the whole exercise exists for. Build a per-case summary first: for each case, how many runs met the expected property, and how the failures differed from each other.

from collections import defaultdict

by_case = defaultdict(list)
for row in rows:
    by_case[row["case_id"]].append(row["output"])

for case in CASES:
    outputs = by_case[case["id"]]
    passes = sum(1 for o in outputs if o == case["expected"])
    print(f"{case['id']}: {passes}/{len(outputs)} exact matches")
    for o in outputs:
        print(f"  - {o}")

Running this against the fixture above produces:

extract_email: 5/5 exact matches
  - ana@example.com
  - ana@example.com
  - ana@example.com
  - ana@example.com
  - ana@example.com
classify_sentiment: 4/5 exact matches
  - negative
  - negative
  - negative
  - mixed
  - negative
format_json: 5/5 exact matches
  - {"name": "Mira", "age": 34}
  - {"name": "Mira", "age": 34}
  - {"name": "Mira", "age": 34}
  - {"name": "Mira", "age": 34}
  - {"name": "Mira", "age": 34}
extract_date: 3/5 exact matches
  - March 3, 2026
  - 03/03/2026
  - March 3, 2026
  - 2026-03-03
  - March 3, 2026

Now classify each case into one of four shapes:

ShapeWhat it looks likeNext action
Stable and correctEvery run passesLeave it alone
Stable and wrongEvery run fails the same wayCapability limit, not a repeatability problem
Unstable but mostly correctPasses most runs, fails someInvestigate the case or the constraint
Unstable and wrongFails often, differently each timeStrong signal of a real weakness

In the fixture, extract_email and format_json are stable and correct. classify_sentiment is unstable but mostly correct — one run drifted to "mixed." extract_date is unstable and mostly correct in meaning but unstable in format: three runs returned the expected string, two returned a different but semantically valid date format. That is surface variation, not semantic failure — and it is exactly the kind of case that a naive exact-match grader would flag as broken.

Now look across cases. A single unstable case is a case problem. Several unstable cases sharing a task type, format, or input length is a task-level pattern. Several cases failing the same way regardless of run is a capability limit — and no amount of repetition will fix it.

A compact grid makes the distinction visible. Mark each cell pass or fail:

            run0  run1  run2  run3  run4
extract_email   P     P     P     P     P
classify_sentiment P  P     P     F     P   <- column of noise
format_json     P     P     P     P     P
extract_date    P     F     P     F     P   <- format drift, not meaning drift

classify_sentiment wobbles; that is within-case variation. extract_date fails on format but not meaning — a different problem that requires a different fix. Averaging across all cases would hide both. Do not average away the interesting case.

Knowledge check

Check your understanding

Answer this question before you continue.

The date case returns the expected date string three times and two different formats for the same date twice. Which interpretation best matches the article?
Scenario Interpretation

Focus: Separate surface-format variation from semantic failure when interpreting case-level results.

What These Observations Do and Do Not Support

A small fixture earns you narrow, honest claims.

What you can say: this case varied across these runs under these parameters; this task type showed repeated failure; this change is worth investigating.

What you cannot say: the model is unreliable in general; the system is fixed; the improvement is real. Small run counts give wide uncertainty, and a non-representative case set gives confident answers about the wrong thing.

There is a subtler trap. If your pass/fail judgment is itself subjective, disagreement between two human graders can look exactly like model instability. Before you blame the model, check whether you would grade the same output the same way twice.

Common mistake: Treating one lucky run as the system's behavior, or one unlucky run as proof the system is broken. Both are single samples dressed up as conclusions.

Extend the Experiment: More Runs, More Cases, One Change

A toy fixture becomes a decision tool when you extend it deliberately.

Add runs to the unstable cases only. More runs give you more observations about that case. They do not resolve whether you are seeing noise or a real distribution — they just narrow the picture. If classify_sentiment keeps producing one "mixed" in every five runs, you have a pattern worth investigating. If the failure rate jumps around unpredictably, you need more runs before you can say anything. Either way, the added runs are evidence to collect, not a verdict to declare.

Add cases that probe the suspected pattern, not more of the same. If format-heavy cases wobble, add more format-heavy cases. If long inputs fail, add long inputs.

Change exactly one thing at a time — a stricter output constraint, a lower temperature, a clearer instruction — and rerun the same fixture. Compare per-case pass counts, not overall averages. A change that fixes classify_sentiment while breaking extract_date is not an improvement; it is a trade you should see clearly.

Record the change and the parameters with the results, so a later reader can tell what was actually tested.

Knowledge check

Check your understanding

Answer this question before you continue.

You suspect an output-format weakness and want to test a stricter format constraint. Which follow-up best supports interpreting the result?
Comparison Reasoning

Focus: Choose an experiment extension that isolates the effect of a system change while retaining case-level evidence.

When to Run This Test, and When Not To

Run it when output stability matters to the user: structured extraction, classification, routing, or any case where the same input should produce the same decision. Run it before blaming the model for a failure you have only seen once.

Skip it when the task is genuinely open-ended and variation is the product — brainstorming, drafting, exploration. There you should measure quality, not agreement. And skip it when you cannot define a checkable expected property, because you will end up measuring your own mood rather than the system.

The Decision Rule

Before you change a prompt, a model, or a parameter, run the same fixed cases several times and read the per-case pattern first. The pattern tells you what kind of problem you have: noise, a case problem, or a task-level weakness.

Your concrete next move: take the fixture above, swap in a real model client, and rerun it. Then pick one case that wobbled, add runs to it, and decide whether you are looking at noise, a case problem, or a task-level weakness. Turn the fixture into a standing regression check, so the next change has to prove itself against the same cases instead of a fresh demo.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

A small fixture shows no variation on its cases. What conclusion is justified?
Question 1 of 2Misconception Check

Focus: State the limits of claims supported by a small number of runs on a small case set.

A fixed-input routing task fails once, and you are considering changing the prompt. What is the best next step according to the article?
Question 2 of 2Scenario Interpretation

Focus: Select an evidence-gathering step before changing a system after a single observed failure.

References

  1. A statistical framework for evaluating the repeatability and reproducibility of large language modelspmc.ncbi.nlm.nih.gov
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.