Build a Small LLM Evaluation Harness
You changed one line in the prompt, ran three examples, and they looked fine. So you shipped. Two weeks later a customer finds the case you never re-ran —…

Key topics
You changed one line in the prompt, ran three examples, and they looked fine. So you shipped. Two weeks later a customer finds the case you never re-ran — and now you are debugging from memory instead of from evidence.
A harness is the smallest piece of infrastructure that turns "it looked good" into a number you can compare. Not a platform. Not a dashboard. A runner, some checks, and a report you can diff. You can build the working version in roughly sixty lines, and that version will catch more regressions than any demo session ever will.
This assumes you already have a small case set with expected criteria. If you do not, start there first — case design is its own skill, and a harness running bad cases just produces confident nonsense faster.
What a Harness Actually Is (and Is Not)
A harness is the runner: it loads fixed cases, calls your system, records raw outputs, applies checks, and writes a comparable report. That is the whole job.
A benchmark is a dataset plus a metric. The harness is the machinery that executes it. Same idea, different scope. When people say "we ran the benchmark," they mean a harness ran it.
Two boundaries worth holding onto:
- Offline evaluation is not a live guardrail. A harness runs against fixed cases, after the fact, to measure behavior. A guardrail runs in the live request path and acts — blocks, retries, escalates. The scoring logic can be identical; the location changes everything.
- A harness is not an observability dashboard. Observability tells you what happened in a real run. A harness tells you whether a change helped, under conditions you control.
Why a script instead of a spreadsheet? Reruns are cheap, diffs are visible, and the same command works before and after a change. A spreadsheet makes you the runner. A script does not.
Knowledge check
Check your understanding
Answer this question before you continue.
The Four Moving Parts
Before writing anything, see the whole system:
cases.jsonl → runner → checks → report.jsonl
↑
system under test
(swappable box)
- Cases: fixed inputs plus expected behavior, stored as data, not buried in code.
- Runner: one function that takes a case and returns the system's output plus metadata you will need later.
- Checks: layered scoring — cheap deterministic checks first, softer judgment checks only where they earn their cost.
- Report: a stable, diffable artifact keyed by case id, so two runs compare line by line.
The system under test is a swappable box. That is the point. You want to change the prompt, the model, or the retrieval step and rerun the same command.
Set Up the Smallest Runnable Version
Dependencies: Python, a JSONL case file, and whatever client your app already uses. No framework required.
Your case file shape:
{"id": "refund-policy-01", "input": "Can I get a refund after 40 days?", "expected": {"must_include": ["30 days"], "must_not_include": ["yes, anytime"]}, "tag": "policy-boundary"}
{"id": "json-format-03", "input": "Return the order as JSON.", "expected": {"valid_json": true, "required_keys": ["order_id", "status"]}, "tag": "format"}
The runner loop, with checks wired in and results persisted:
import json
from your_app import answer # your system under test
def check(case, output):
scores = {}
exp = case["expected"]
if "must_include" in exp:
scores["contains"] = all(s in output for s in exp["must_include"])
if "must_not_include" in exp:
scores["excludes"] = not any(s in output for s in exp["must_not_include"])
if exp.get("valid_json"):
try:
parsed = json.loads(output)
scores["valid_json"] = True
scores["keys"] = all(k in parsed for k in exp.get("required_keys", []))
except json.JSONDecodeError:
scores["valid_json"] = False
return scores
with open("cases.jsonl") as f:
cases = [json.loads(line) for line in f]
results = []
for case in cases:
output = answer(case["input"])
scores = check(case, output)
results.append({
"id": case["id"],
"tag": case.get("tag"),
"input": case["input"],
"expected": case["expected"],
"output": output,
"scores": scores,
"passed": all(scores.values()) if scores else None,
"system_version": "prompt-v3",
})
with open("results.jsonl", "w") as f:
for r in results:
f.write(json.dumps(r) + "\n")
print(f"ran {len(results)} cases")
Expected output: a results.jsonl with one row per case, each carrying its raw output, its per-check scores, and a passed flag. The printed count is your first observable success criterion. If it says ran 12 cases and you have 12 cases, the runner works.
Store raw outputs before scoring. You will re-score old runs later, and you cannot re-score what you did not save. This one decision separates a harness from a throwaway script.
Knowledge check
Check your understanding
Answer this question before you continue.
Score in Layers, Not in One Number
Layer one: deterministic checks. Required substrings, valid JSON, schema conformance, refusal detection, length bounds. Fast, free, unambiguous.
Layer two: rubric or model-based judgment for open-ended quality, used only on cases where deterministic checks cannot decide.
The order matters because a failing deterministic check is a fact, while a failing judge score is an opinion that needs calibration. Do not let an opinion override a fact.
Common mistake: Collapsing everything into a single pass rate. It hides which behavior broke and makes regressions invisible. Keep the score per case and per check. The aggregate is for the report; the per-case detail is for the diagnosis.
Knowledge check
Check your understanding
Answer this question before you continue.
Label Failures While the Evidence Is Fresh
A failing run is not a verdict. It is a dataset waiting to be labeled.
Label categories that actually drive action:
| Label | What it means | Typical fix |
|---|---|---|
| Wrong answer | Output contradicts expected behavior | Prompt or model change |
| Missing evidence | Answer is plausible but unsupported | Retrieval or grounding |
| Format violation | Shape is wrong, content may be fine | Output parsing or schema |
| Refusal | System declined a valid request | Prompt or safety tuning |
| Hallucinated detail | Invented specifics | Grounding, temperature, retrieval |
| Correct-but-unhelpful | True but useless to the user | Prompt framing |
Label the case, not the run. A case that fails for two different reasons across runs is telling you something about instability — that is a finding, not noise.
Store the label next to the raw output so the next person sees the evidence, not just the verdict.
Failure mode to expect: The first run produces labels you did not anticipate. That is the harness working, not the harness failing. If a label does not change what you would fix, it is not a useful label — cut it.
Run the Before/After Comparison
This is the payoff. Freeze everything except the change: same cases, same checks, same model version, same parameters. Otherwise you are measuring the environment, not the change.
Diff by case id, not by aggregate score:
def load(path):
with open(path) as f:
return [json.loads(line) for line in f]
before = {r["id"]: r for r in load("results_before.jsonl")}
after = {r["id"]: r for r in load("results_after.jsonl")}
fixed, broke = [], []
for cid in before:
b = before[cid]["passed"]
a = after[cid]["passed"]
if not b and a: fixed.append(cid)
if b and not a: broke.append(cid)
print("fixed:", fixed)
print("broke:", broke)
The interesting output is the list of cases that flipped in each direction. Read the trade: a change that fixes the targeted behavior while breaking three unrelated cases is a net loss you would have missed from demos.
Small-sample honesty: With a few dozen cases, a one-case difference is noise. Say what the run supports and what it does not. A harness that produces confident wrong conclusions is worse than no harness.
Decision rule: Ship the change when the targeted cases improve and no previously passing case regresses in a way that matters. A single flipped case is not an automatic veto — inspect it. If the flip is a real regression on a behavior you care about, block the ship. If it is a borderline case where the new output is arguably better, note it, adjust the expected criteria if needed, and re-run. The rule is not "zero flips." The rule is "no unexplained flips."
Knowledge check
Check your understanding
Answer this question before you continue.
Make It Repeatable, Then Make It Bigger
One meaningful experiment: add a new case drawn from a real failure you have seen, rerun, and confirm the harness catches it. If it does not, your checks are too weak — that is the harness telling you where to grow.
Then wire the same command into a pre-merge check so the comparison runs without anyone remembering to run it.
When to grow:
- Add cases when you find a real failure.
- Add checks when a label keeps recurring.
- Add judge calibration when scores disagree with your reading.
When not to grow: Do not build a distributed evaluation platform for a feature with thirty cases and one owner. Cost and time are real constraints — run cheap deterministic checks on every change and reserve expensive judgment checks for the changes that matter.
The harness tells you that something broke. It does not always tell you why. When a case flips, the next move is to diagnose which layer actually caused it — prompt, model, retrieval, or workflow — and that is a separate skill worth building next.
Knowledge check
Final check
Finish the article by checking the ideas you just learned.
Want a more structured LLMOps path?
Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.
Large Language Models Starter Pack
A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.
- 227-page Illustrated PDF edition
- 12 guided LLM engineering chapters
- Visual concept diagrams
- Self-assessment quizzes
- Bonus deep-dive sections
- Prompt design, structured output, context windows & RAG pipelines
- Agents, tool calling, prompt injection, evaluation & application lifecycles
Coming soon


