Skip to content
intermediate

Build a Small LLM Evaluation Harness

You changed one line in the prompt, ran three examples, and they looked fine. So you shipped. Two weeks later a customer finds the case you never re-ran —…

Published 2026-10-03Updated 2026-10-048 min read
A detailed view of rippling sand patterns in a desert, showcasing natural textures.
A detailed view of rippling sand patterns in a desert, showcasing natural textures. Photo by Edoardo Tommasini on Pexels.

You changed one line in the prompt, ran three examples, and they looked fine. So you shipped. Two weeks later a customer finds the case you never re-ran — and now you are debugging from memory instead of from evidence.

A harness is the smallest piece of infrastructure that turns "it looked good" into a number you can compare. Not a platform. Not a dashboard. A runner, some checks, and a report you can diff. You can build the working version in roughly sixty lines, and that version will catch more regressions than any demo session ever will.

This assumes you already have a small case set with expected criteria. If you do not, start there first — case design is its own skill, and a harness running bad cases just produces confident nonsense faster.

What a Harness Actually Is (and Is Not)

A harness is the runner: it loads fixed cases, calls your system, records raw outputs, applies checks, and writes a comparable report. That is the whole job.

A benchmark is a dataset plus a metric. The harness is the machinery that executes it. Same idea, different scope. When people say "we ran the benchmark," they mean a harness ran it.

Two boundaries worth holding onto:

  • Offline evaluation is not a live guardrail. A harness runs against fixed cases, after the fact, to measure behavior. A guardrail runs in the live request path and acts — blocks, retries, escalates. The scoring logic can be identical; the location changes everything.
  • A harness is not an observability dashboard. Observability tells you what happened in a real run. A harness tells you whether a change helped, under conditions you control.

Why a script instead of a spreadsheet? Reruns are cheap, diffs are visible, and the same command works before and after a change. A spreadsheet makes you the runner. A script does not.

Knowledge check

Check your understanding

Answer this question before you continue.

A team reruns fixed cases after each prompt edit, records outputs, and compares reports. Which description best fits this tool?
Scenario Interpretation

Focus: Distinguish an offline evaluation harness from a live guardrail based on where and how it operates.

The Four Moving Parts

Before writing anything, see the whole system:

cases.jsonl  →  runner  →  checks  →  report.jsonl
                  ↑
          system under test
          (swappable box)
  • Cases: fixed inputs plus expected behavior, stored as data, not buried in code.
  • Runner: one function that takes a case and returns the system's output plus metadata you will need later.
  • Checks: layered scoring — cheap deterministic checks first, softer judgment checks only where they earn their cost.
  • Report: a stable, diffable artifact keyed by case id, so two runs compare line by line.

The system under test is a swappable box. That is the point. You want to change the prompt, the model, or the retrieval step and rerun the same command.

Set Up the Smallest Runnable Version

Dependencies: Python, a JSONL case file, and whatever client your app already uses. No framework required.

Your case file shape:

{"id": "refund-policy-01", "input": "Can I get a refund after 40 days?", "expected": {"must_include": ["30 days"], "must_not_include": ["yes, anytime"]}, "tag": "policy-boundary"}
{"id": "json-format-03", "input": "Return the order as JSON.", "expected": {"valid_json": true, "required_keys": ["order_id", "status"]}, "tag": "format"}

The runner loop, with checks wired in and results persisted:

import json
from your_app import answer  # your system under test

def check(case, output):
    scores = {}
    exp = case["expected"]
    if "must_include" in exp:
        scores["contains"] = all(s in output for s in exp["must_include"])
    if "must_not_include" in exp:
        scores["excludes"] = not any(s in output for s in exp["must_not_include"])
    if exp.get("valid_json"):
        try:
            parsed = json.loads(output)
            scores["valid_json"] = True
            scores["keys"] = all(k in parsed for k in exp.get("required_keys", []))
        except json.JSONDecodeError:
            scores["valid_json"] = False
    return scores

with open("cases.jsonl") as f:
    cases = [json.loads(line) for line in f]

results = []
for case in cases:
    output = answer(case["input"])
    scores = check(case, output)
    results.append({
        "id": case["id"],
        "tag": case.get("tag"),
        "input": case["input"],
        "expected": case["expected"],
        "output": output,
        "scores": scores,
        "passed": all(scores.values()) if scores else None,
        "system_version": "prompt-v3",
    })

with open("results.jsonl", "w") as f:
    for r in results:
        f.write(json.dumps(r) + "\n")

print(f"ran {len(results)} cases")

Expected output: a results.jsonl with one row per case, each carrying its raw output, its per-check scores, and a passed flag. The printed count is your first observable success criterion. If it says ran 12 cases and you have 12 cases, the runner works.

Store raw outputs before scoring. You will re-score old runs later, and you cannot re-score what you did not save. This one decision separates a harness from a throwaway script.

Knowledge check

Check your understanding

Answer this question before you continue.

A harness stores only each case's pass/fail result. What important capability has this design lost?
Debugging

Focus: Explain why a harness should persist raw outputs before scoring them.

Score in Layers, Not in One Number

Layer one: deterministic checks. Required substrings, valid JSON, schema conformance, refusal detection, length bounds. Fast, free, unambiguous.

Layer two: rubric or model-based judgment for open-ended quality, used only on cases where deterministic checks cannot decide.

The order matters because a failing deterministic check is a fact, while a failing judge score is an opinion that needs calibration. Do not let an opinion override a fact.

Common mistake: Collapsing everything into a single pass rate. It hides which behavior broke and makes regressions invisible. Keep the score per case and per check. The aggregate is for the report; the per-case detail is for the diagnosis.

Knowledge check

Check your understanding

Answer this question before you continue.

A report shows only an overall pass rate, which falls after a change. What is the most useful improvement to make the regression diagnosable?
Misconception Check

Focus: Use per-case and per-check scores to preserve diagnostic information rather than relying only on an aggregate pass rate.

Label Failures While the Evidence Is Fresh

A failing run is not a verdict. It is a dataset waiting to be labeled.

Label categories that actually drive action:

LabelWhat it meansTypical fix
Wrong answerOutput contradicts expected behaviorPrompt or model change
Missing evidenceAnswer is plausible but unsupportedRetrieval or grounding
Format violationShape is wrong, content may be fineOutput parsing or schema
RefusalSystem declined a valid requestPrompt or safety tuning
Hallucinated detailInvented specificsGrounding, temperature, retrieval
Correct-but-unhelpfulTrue but useless to the userPrompt framing

Label the case, not the run. A case that fails for two different reasons across runs is telling you something about instability — that is a finding, not noise.

Store the label next to the raw output so the next person sees the evidence, not just the verdict.

Failure mode to expect: The first run produces labels you did not anticipate. That is the harness working, not the harness failing. If a label does not change what you would fix, it is not a useful label — cut it.

Run the Before/After Comparison

The same cases and checks feed two parallel runs: a baseline system and a changed system. Their reports meet at a case-ID diff that separates fixed cases from broken cases.
Keep the test conditions fixed, then inspect which individual cases improved or regressed.

This is the payoff. Freeze everything except the change: same cases, same checks, same model version, same parameters. Otherwise you are measuring the environment, not the change.

Diff by case id, not by aggregate score:

def load(path):
    with open(path) as f:
        return [json.loads(line) for line in f]

before = {r["id"]: r for r in load("results_before.jsonl")}
after  = {r["id"]: r for r in load("results_after.jsonl")}

fixed, broke = [], []
for cid in before:
    b = before[cid]["passed"]
    a = after[cid]["passed"]
    if not b and a: fixed.append(cid)
    if b and not a: broke.append(cid)

print("fixed:", fixed)
print("broke:", broke)

The interesting output is the list of cases that flipped in each direction. Read the trade: a change that fixes the targeted behavior while breaking three unrelated cases is a net loss you would have missed from demos.

Small-sample honesty: With a few dozen cases, a one-case difference is noise. Say what the run supports and what it does not. A harness that produces confident wrong conclusions is worse than no harness.

Decision rule: Ship the change when the targeted cases improve and no previously passing case regresses in a way that matters. A single flipped case is not an automatic veto — inspect it. If the flip is a real regression on a behavior you care about, block the ship. If it is a borderline case where the new output is arguably better, note it, adjust the expected criteria if needed, and re-run. The rule is not "zero flips." The rule is "no unexplained flips."

Knowledge check

Check your understanding

Answer this question before you continue.

A team wants to know whether a prompt edit improved its targeted behavior. Which comparison best isolates the effect of that edit?
Comparison Reasoning

Focus: Design a before/after comparison that attributes observed changes to the system change being evaluated.

Make It Repeatable, Then Make It Bigger

One meaningful experiment: add a new case drawn from a real failure you have seen, rerun, and confirm the harness catches it. If it does not, your checks are too weak — that is the harness telling you where to grow.

Then wire the same command into a pre-merge check so the comparison runs without anyone remembering to run it.

When to grow:

  • Add cases when you find a real failure.
  • Add checks when a label keeps recurring.
  • Add judge calibration when scores disagree with your reading.

When not to grow: Do not build a distributed evaluation platform for a feature with thirty cases and one owner. Cost and time are real constraints — run cheap deterministic checks on every change and reserve expensive judgment checks for the changes that matter.

The harness tells you that something broke. It does not always tell you why. When a case flips, the next move is to diagnose which layer actually caused it — prompt, model, retrieval, or workflow — and that is a separate skill worth building next.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

The same case fails in two runs, once with a format violation and once with a hallucinated detail. How should the team interpret and record this?
Question 1 of 2Scenario Interpretation

Focus: Label failures at the case level and recognize different failure reasons across runs as evidence of instability.

A team has a small harness and encounters a real production failure not covered by its cases. What is the best next step?
Question 2 of 2Scenario Interpretation

Focus: Choose a proportionate next step for extending a harness and confirm that a new case can detect a real failure.

Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.