Practice Auditing an LLM Evaluation Set for Leakage and Contamination
Your eval score went up. The change shipped. Nobody asked whether the test set was ever clean.

Key topics
Your eval score went up. The change shipped. Nobody asked whether the test set was ever clean.
That is the uncomfortable part of evaluation work: a score is a measurement taken under test conditions, and if the conditions are broken, the number is not evidence — it is decoration. This exercise is a hygiene pass on the evidence itself. You will audit a small synthetic fixture for three failure families, write a ledger that names what each finding compromises, and run a checker that tells you what you missed.
The deliverable is a ledger, not a vibe.
Why a Passing Score Can Still Be Wrong
An evaluation result is a claim about a fixture. "Accuracy improved six points" really means "on these cases, under these conditions, the system produced more acceptable outputs." Break the conditions and the claim collapses — not because the model got worse, but because the test stopped testing what you think it tested.
Three contamination families cause most of this damage:
- Exact and near duplicates. The same underlying case appears twice, so a small fixture looks broader than it is and one fix looks like a general improvement.
- Answer-key or prompt leakage. The expected answer, or a near-verbatim fragment of it, sits inside the model's input. The test becomes a lookup.
- Split contamination. An eval case, or a close variant, already appears in your development history. You are grading the system on something it was tuned on.
One boundary matters before you start. This is not the same problem as a model memorizing its pretraining data. That leak lives inside weights you cannot inspect. Here the leak lives inside your own JSONL files — which is exactly why you can fix it.
And a finding without an affected claim is trivia. Every row in your ledger has to name the specific claim it damages, or it does not belong there.
Set Up the Fixture and Inspect the Records
Environment assumption: Python 3 standard library only. No installs, no API keys, no network. The modules you need are json, collections, difflib, hashlib, and argparse.
You have three inputs:
eval.jsonl— the test casesdev_history.jsonl— development and prompt historyanswers.json— the ledger you will submit
Start by parsing and printing the shape of the data. You want to see the records before you judge them.
import json
def load_jsonl(path):
records = []
with open(path, encoding="utf-8") as f:
for lineno, line in enumerate(f, start=1):
line = line.strip()
if not line:
continue
try:
records.append(json.loads(line))
except json.JSONDecodeError as e:
raise ValueError(f"{path}:{lineno} malformed JSON: {e}") from e
return records
evals = load_jsonl("eval.jsonl")
dev = load_jsonl("dev_history.jsonl")
print(f"eval records: {len(evals)} dev records: {len(dev)}")
for r in evals[:3]:
print(r["id"], "|", r["prompt"][:70], "|", r["expected"][:40])
Expected output: a stable record count and a readable dump. Two debugging signals to watch for:
- A JSONL parsing error means a fixture line is malformed. That is a signal about the fixture, not a reason to skip the record.
- A count that changes between runs means your parser is silently dropping lines. Stop and fix the parser before you trust anything downstream.
Knowledge check
Check your understanding
Answer this question before you continue.
Find Exact and Near Duplicates
Exact duplicates are cheap and deterministic. Normalize whitespace and case, hash the prompt, group by hash.
import hashlib, re
from collections import defaultdict
def norm(text):
return re.sub(r"\s+", " ", text.strip().lower())
def h(text):
return hashlib.sha256(norm(text).encode()).hexdigest()
by_hash = defaultdict(list)
for r in evals:
by_hash[h(r["prompt"])].append(r["id"])
for digest, ids in by_hash.items():
if len(ids) > 1:
print("exact duplicate:", ids)
Near duplicates need a similarity score. difflib.SequenceMatcher is slow but dependency-free and good enough for a fixture this size.
from difflib import SequenceMatcher
def ratio(a, b):
return SequenceMatcher(None, norm(a), norm(b)).ratio()
for i in range(len(evals)):
for j in range(i + 1, len(evals)):
score = ratio(evals[i]["prompt"], evals[j]["prompt"])
if score >= 0.85:
print(f"{evals[i]['id']} ~ {evals[j]['id']} {score:.2f}")
Why near-duplicate test cases inflate results: the same underlying case counted twice makes a small fixture look broader than it is, and a single fix looks like a general improvement. If three of your "improved" cases are one case wearing three hats, your six-point gain is closer to a two-point gain.
The threshold is a judgment, not a fact. Report the cutoff you used and the pairs sitting just below it, so a reviewer can disagree with your line.
Common mistake: deduplicating on the prompt while ignoring that two different prompts share one expected answer. That is a coverage gap, not a duplicate — and collapsing them would hide it.
Knowledge check
Check your understanding
Answer this question before you continue.
Detect Answer-Key and Prompt Leakage
This is the finding most likely to invalidate a headline claim outright. If the graded answer is sitting in the input, the model is not reasoning. It is reading.
Check containment in both directions: between each eval prompt and its own expected answer, and between eval prompts and dev history entries.
def tokens(text):
return set(norm(text).split())
def containment(needle, haystack):
n, hset = tokens(needle), tokens(haystack)
if not n:
return 0.0
return len(n & hset) / len(n)
for r in evals:
score = containment(r["expected"], r["prompt"])
if score > 0.6:
print(f"{r['id']} answer leakage: {score:.2f} of expected tokens in prompt")
for r in evals:
for d in dev:
score = containment(r["prompt"], d.get("prompt", ""))
if score > 0.8:
print(f"{r['id']} prompt seen in dev: {d['id']} {score:.2f}")
Report the matched span, not just a boolean. The span is what makes the finding reviewable and what tells you how much of the answer was given away.
Note: A worked example in the prompt is a design choice. The graded answer sitting in the prompt is a broken test. The difference is whether the leaked text is the thing being scored.
Knowledge check
Check your understanding
Answer this question before you continue.
Check Split Contamination Between Dev History and Eval
Split contamination is the quietest failure. An eval case, or a close variant, already appears in dev_history.jsonl, so the system was tuned on something it is now being graded on.
Reuse the duplicate detector across files rather than within one file, and record which side each match came from.
for e in evals:
for d in dev:
score = ratio(e["prompt"], d.get("prompt", ""))
if score >= 0.85:
print(f"contamination: eval {e['id']} ~ dev {d['id']} {score:.2f}")
Why this matters for claims: a before/after improvement on a contaminated case measures recall of a seen example, not generalization. The number is real. The conclusion you drew from it is not.
Warning: Unmatched IDs are a debugging signal. An eval id with no counterpart in your ledger usually means a parsing or join bug, not a clean result. Chase it before you celebrate.
One boundary: overlap is not automatically fatal. A deliberately shared smoke case is fine if it is labeled and excluded from the headline metric. The failure is unlabeled overlap, not overlap itself.
Knowledge check
Check your understanding
Answer this question before you continue.
Write the Audit Ledger
Detection is not the deliverable. The ledger is. This is where most audits quietly fail — they stop at a list of problems and never convert findings into decisions.
Keep one finding per row, with stable ids and no prose paragraphs inside fields:
| Field | Meaning |
|---|---|
finding_id | Stable identifier, e.g. F-001 |
type | duplicate, leakage, or contamination |
evidence | Record ids and the matched span |
affected_claim | The specific claim this compromises |
severity | How much of the claim it damages |
action | remediate or qualify |
The affected claim column is load-bearing. "Eval accuracy improved six points" is compromised if three of the improved cases are duplicates. "The system handles edge cases" is compromised if the edge cases leaked into the prompt.
Two bounded actions:
- Remediate — remove, rewrite, or re-split the case, then re-run.
- Qualify — keep the case, weaken the claim, and say so in the write-up.
Prefer qualification over silent deletion when the case is genuinely valuable. Deleting evidence to protect a number is the exact failure this exercise exists to prevent.
Shape of answers.json
The checker reads a JSON file, not a Markdown table. Keep the same fields, one object per finding, wrapped in a top-level list. Here is a minimal placeholder row — replace the values with your own findings; the ids and spans below are illustrative, not the fixture's answers.
[
{
"finding_id": "F-001",
"type": "duplicate",
"evidence": "eval ids e-004 and e-011, prompt similarity 0.91",
"affected_claim": "accuracy improved on the edge-case slice",
"severity": "high",
"action": "remediate"
},
{
"finding_id": "F-002",
"type": "leakage",
"evidence": "eval id e-007, 0.72 of expected tokens appear in prompt",
"affected_claim": "the model reasons about the task rather than copying",
"severity": "high",
"action": "qualify"
}
]
Two rules keep the submission valid:
typemust be one ofduplicate,leakage, orcontamination. Anything else is a schema problem, not a finding.actionmust beremediateorqualify. Free-text actions will not match.
If the checker rejects your file before evaluating findings, the problem is shape, not substance. Fix the JSON first, then re-read the feedback.
Run the Checker and Read the Feedback
Now get an external verdict instead of grading yourself.
python audit_eval.py eval.jsonl dev_history.jsonl answers.json
Read the output as a diff against the documented injected issues. A clean run means your ledger named them. A partial run tells you which family you missed — which is more useful than a score.
Expect three debugging signals, and treat them differently:
- JSONL parsing errors point at malformed fixture lines.
- Schema errors — unknown
type, missing field, malformed JSON inanswers.json— mean the checker could not read your submission. Fix the file; do not rewrite findings. - Unmatched IDs mean the checker read your file but could not join a finding to a fixture record. That is usually an id-normalization bug in your ledger, not a missed finding.
Tip: Do not tune the ledger to the checker. If you disagree with a flagged finding, record the disagreement and the evidence. That is a legitimate audit outcome, not a failure.
Alter a Record and Audit Again
One fixture is not a procedure. Prove the audit generalizes by mutating a record and predicting the consequence before you run.
Pick a mutation with a predictable outcome:
- Copy an eval prompt into
dev_history.jsonlto inject split contamination. - Paste the expected answer into a prompt to inject answer leakage.
State your prediction, re-run the checker, and confirm the new finding appears with the right affected claim. If your prediction and the output disagree, the disagreement is the lesson.
Then try a mutation that should not be caught: a paraphrase that shares meaning but not tokens. Note the limit honestly. Token and substring detection misses semantic duplicates and paraphrased leakage, and a clean run is not proof of cleanliness.
What This Audit Does Not Prove
A clean ledger says the fixture is internally consistent. It says nothing about whether the cases are representative of real inputs. Those are different questions with different tools.
Three boundaries to hold:
- Token and substring detection misses semantic duplicates and paraphrased leakage.
- Contamination inside a model's pretraining data is a different problem with different tooling. Auditing your own JSONL files cannot resolve it.
- A clean audit is a necessary condition for trusting a score, not a sufficient one.
The honest output of a clean audit is a qualified claim: on this fixture, after this audit, the improvement holds for these cases.
The Decision Rule
Before you quote an eval number, run the audit and attach the ledger. The number travels with its evidence, or it does not travel.
Your next move is not a new fixture. Audit the one you already have. Run the checker, then mutate one record and run it again, so the procedure becomes yours rather than this article's. When you can predict what a mutation will produce before the checker confirms it, you have stopped memorizing a fixture and started owning a method.
Knowledge check
Final check
Finish the article by checking the ideas you just learned.
References
Want a more structured LLMOps path?
Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.
Large Language Models Starter Pack
A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.
- 227-page Illustrated PDF edition
- 12 guided LLM engineering chapters
- Visual concept diagrams
- Self-assessment quizzes
- Bonus deep-dive sections
- Prompt design, structured output, context windows & RAG pipelines
- Agents, tool calling, prompt injection, evaluation & application lifecycles
Coming soon


