Skip to content
intermediate

Practice Auditing an LLM Evaluation Set for Leakage and Contamination

Your eval score went up. The change shipped. Nobody asked whether the test set was ever clean.

Published 2026-10-03Updated 2026-10-0410 min read
Captivating view of sand dunes and landscape at Skagen, Denmark, under a cloudy sky.
Captivating view of sand dunes and landscape at Skagen, Denmark, under a cloudy sky. Photo by Tommes Frites on Pexels.

Your eval score went up. The change shipped. Nobody asked whether the test set was ever clean.

That is the uncomfortable part of evaluation work: a score is a measurement taken under test conditions, and if the conditions are broken, the number is not evidence — it is decoration. This exercise is a hygiene pass on the evidence itself. You will audit a small synthetic fixture for three failure families, write a ledger that names what each finding compromises, and run a checker that tells you what you missed.

The deliverable is a ledger, not a vibe.

Why a Passing Score Can Still Be Wrong

An evaluation claim is connected to three failure modes: duplicates inflate apparent coverage, answer leakage turns a test into lookup, and dev-history overlap tests seen cases. Each weakens the claim that the score demonstrates generalization.
Different fixture defects weaken different parts of an evaluation claim; identify the failure before deciding what the score supports.

An evaluation result is a claim about a fixture. "Accuracy improved six points" really means "on these cases, under these conditions, the system produced more acceptable outputs." Break the conditions and the claim collapses — not because the model got worse, but because the test stopped testing what you think it tested.

Three contamination families cause most of this damage:

  • Exact and near duplicates. The same underlying case appears twice, so a small fixture looks broader than it is and one fix looks like a general improvement.
  • Answer-key or prompt leakage. The expected answer, or a near-verbatim fragment of it, sits inside the model's input. The test becomes a lookup.
  • Split contamination. An eval case, or a close variant, already appears in your development history. You are grading the system on something it was tuned on.

One boundary matters before you start. This is not the same problem as a model memorizing its pretraining data. That leak lives inside weights you cannot inspect. Here the leak lives inside your own JSONL files — which is exactly why you can fix it.

And a finding without an affected claim is trivia. Every row in your ledger has to name the specific claim it damages, or it does not belong there.

Set Up the Fixture and Inspect the Records

Environment assumption: Python 3 standard library only. No installs, no API keys, no network. The modules you need are json, collections, difflib, hashlib, and argparse.

You have three inputs:

  • eval.jsonl — the test cases
  • dev_history.jsonl — development and prompt history
  • answers.json — the ledger you will submit

Start by parsing and printing the shape of the data. You want to see the records before you judge them.

import json

def load_jsonl(path):
    records = []
    with open(path, encoding="utf-8") as f:
        for lineno, line in enumerate(f, start=1):
            line = line.strip()
            if not line:
                continue
            try:
                records.append(json.loads(line))
            except json.JSONDecodeError as e:
                raise ValueError(f"{path}:{lineno} malformed JSON: {e}") from e
    return records

evals = load_jsonl("eval.jsonl")
dev = load_jsonl("dev_history.jsonl")
print(f"eval records: {len(evals)}  dev records: {len(dev)}")
for r in evals[:3]:
    print(r["id"], "|", r["prompt"][:70], "|", r["expected"][:40])

Expected output: a stable record count and a readable dump. Two debugging signals to watch for:

  • A JSONL parsing error means a fixture line is malformed. That is a signal about the fixture, not a reason to skip the record.
  • A count that changes between runs means your parser is silently dropping lines. Stop and fix the parser before you trust anything downstream.

Knowledge check

Check your understanding

Answer this question before you continue.

Your JSONL inspection script reports a different eval-record count on successive runs. What is the best next step?
Debugging

Focus: Diagnose a changing record count as evidence of a parser problem before interpreting audit results.

Find Exact and Near Duplicates

Exact duplicates are cheap and deterministic. Normalize whitespace and case, hash the prompt, group by hash.

import hashlib, re
from collections import defaultdict

def norm(text):
    return re.sub(r"\s+", " ", text.strip().lower())

def h(text):
    return hashlib.sha256(norm(text).encode()).hexdigest()

by_hash = defaultdict(list)
for r in evals:
    by_hash[h(r["prompt"])].append(r["id"])

for digest, ids in by_hash.items():
    if len(ids) > 1:
        print("exact duplicate:", ids)

Near duplicates need a similarity score. difflib.SequenceMatcher is slow but dependency-free and good enough for a fixture this size.

from difflib import SequenceMatcher

def ratio(a, b):
    return SequenceMatcher(None, norm(a), norm(b)).ratio()

for i in range(len(evals)):
    for j in range(i + 1, len(evals)):
        score = ratio(evals[i]["prompt"], evals[j]["prompt"])
        if score >= 0.85:
            print(f"{evals[i]['id']} ~ {evals[j]['id']}  {score:.2f}")

Why near-duplicate test cases inflate results: the same underlying case counted twice makes a small fixture look broader than it is, and a single fix looks like a general improvement. If three of your "improved" cases are one case wearing three hats, your six-point gain is closer to a two-point gain.

The threshold is a judgment, not a fact. Report the cutoff you used and the pairs sitting just below it, so a reviewer can disagree with your line.

Common mistake: deduplicating on the prompt while ignoring that two different prompts share one expected answer. That is a coverage gap, not a duplicate — and collapsing them would hide it.

Knowledge check

Check your understanding

Answer this question before you continue.

After using a similarity cutoff to flag likely near-duplicate prompts, what should the auditor report to make the judgment reviewable?
Single Choice

Focus: Explain how to make a near-duplicate threshold reviewable rather than treating the cutoff as objective truth.

Detect Answer-Key and Prompt Leakage

This is the finding most likely to invalidate a headline claim outright. If the graded answer is sitting in the input, the model is not reasoning. It is reading.

Check containment in both directions: between each eval prompt and its own expected answer, and between eval prompts and dev history entries.

def tokens(text):
    return set(norm(text).split())

def containment(needle, haystack):
    n, hset = tokens(needle), tokens(haystack)
    if not n:
        return 0.0
    return len(n & hset) / len(n)

for r in evals:
    score = containment(r["expected"], r["prompt"])
    if score > 0.6:
        print(f"{r['id']} answer leakage: {score:.2f} of expected tokens in prompt")

for r in evals:
    for d in dev:
        score = containment(r["prompt"], d.get("prompt", ""))
        if score > 0.8:
            print(f"{r['id']} prompt seen in dev: {d['id']}  {score:.2f}")

Report the matched span, not just a boolean. The span is what makes the finding reviewable and what tells you how much of the answer was given away.

Note: A worked example in the prompt is a design choice. The graded answer sitting in the prompt is a broken test. The difference is whether the leaked text is the thing being scored.

Knowledge check

Check your understanding

Answer this question before you continue.

An eval case has a containment score of 0.72 when comparing its expected answer with its prompt. Which ledger evidence best supports reviewing this suspected leakage?
Scenario Interpretation

Focus: Interpret answer-token containment as evidence of potential answer leakage and identify the evidence to record.

Check Split Contamination Between Dev History and Eval

Split contamination is the quietest failure. An eval case, or a close variant, already appears in dev_history.jsonl, so the system was tuned on something it is now being graded on.

Reuse the duplicate detector across files rather than within one file, and record which side each match came from.

for e in evals:
    for d in dev:
        score = ratio(e["prompt"], d.get("prompt", ""))
        if score >= 0.85:
            print(f"contamination: eval {e['id']} ~ dev {d['id']}  {score:.2f}")

Why this matters for claims: a before/after improvement on a contaminated case measures recall of a seen example, not generalization. The number is real. The conclusion you drew from it is not.

Warning: Unmatched IDs are a debugging signal. An eval id with no counterpart in your ledger usually means a parsing or join bug, not a clean result. Chase it before you celebrate.

One boundary: overlap is not automatically fatal. A deliberately shared smoke case is fine if it is labeled and excluded from the headline metric. The failure is unlabeled overlap, not overlap itself.

Knowledge check

Check your understanding

Answer this question before you continue.

An eval case also appears in development history, but it is deliberately labeled as a shared smoke case and excluded from the headline metric. How should the auditor interpret this overlap?
Scenario Interpretation

Focus: Distinguish disclosed, excluded shared smoke cases from unlabeled split contamination that undermines a headline claim.

Write the Audit Ledger

Detection is not the deliverable. The ledger is. This is where most audits quietly fail — they stop at a list of problems and never convert findings into decisions.

Keep one finding per row, with stable ids and no prose paragraphs inside fields:

FieldMeaning
finding_idStable identifier, e.g. F-001
typeduplicate, leakage, or contamination
evidenceRecord ids and the matched span
affected_claimThe specific claim this compromises
severityHow much of the claim it damages
actionremediate or qualify

The affected claim column is load-bearing. "Eval accuracy improved six points" is compromised if three of the improved cases are duplicates. "The system handles edge cases" is compromised if the edge cases leaked into the prompt.

Two bounded actions:

  • Remediate — remove, rewrite, or re-split the case, then re-run.
  • Qualify — keep the case, weaken the claim, and say so in the write-up.

Prefer qualification over silent deletion when the case is genuinely valuable. Deleting evidence to protect a number is the exact failure this exercise exists to prevent.

Shape of answers.json

The checker reads a JSON file, not a Markdown table. Keep the same fields, one object per finding, wrapped in a top-level list. Here is a minimal placeholder row — replace the values with your own findings; the ids and spans below are illustrative, not the fixture's answers.

[
  {
    "finding_id": "F-001",
    "type": "duplicate",
    "evidence": "eval ids e-004 and e-011, prompt similarity 0.91",
    "affected_claim": "accuracy improved on the edge-case slice",
    "severity": "high",
    "action": "remediate"
  },
  {
    "finding_id": "F-002",
    "type": "leakage",
    "evidence": "eval id e-007, 0.72 of expected tokens appear in prompt",
    "affected_claim": "the model reasons about the task rather than copying",
    "severity": "high",
    "action": "qualify"
  }
]

Two rules keep the submission valid:

  • type must be one of duplicate, leakage, or contamination. Anything else is a schema problem, not a finding.
  • action must be remediate or qualify. Free-text actions will not match.

If the checker rejects your file before evaluating findings, the problem is shape, not substance. Fix the JSON first, then re-read the feedback.

Run the Checker and Read the Feedback

Now get an external verdict instead of grading yourself.

python audit_eval.py eval.jsonl dev_history.jsonl answers.json

Read the output as a diff against the documented injected issues. A clean run means your ledger named them. A partial run tells you which family you missed — which is more useful than a score.

Expect three debugging signals, and treat them differently:

  • JSONL parsing errors point at malformed fixture lines.
  • Schema errors — unknown type, missing field, malformed JSON in answers.json — mean the checker could not read your submission. Fix the file; do not rewrite findings.
  • Unmatched IDs mean the checker read your file but could not join a finding to a fixture record. That is usually an id-normalization bug in your ledger, not a missed finding.

Tip: Do not tune the ledger to the checker. If you disagree with a flagged finding, record the disagreement and the evidence. That is a legitimate audit outcome, not a failure.

Alter a Record and Audit Again

One fixture is not a procedure. Prove the audit generalizes by mutating a record and predicting the consequence before you run.

Pick a mutation with a predictable outcome:

  • Copy an eval prompt into dev_history.jsonl to inject split contamination.
  • Paste the expected answer into a prompt to inject answer leakage.

State your prediction, re-run the checker, and confirm the new finding appears with the right affected claim. If your prediction and the output disagree, the disagreement is the lesson.

Then try a mutation that should not be caught: a paraphrase that shares meaning but not tokens. Note the limit honestly. Token and substring detection misses semantic duplicates and paraphrased leakage, and a clean run is not proof of cleanliness.

What This Audit Does Not Prove

A clean ledger says the fixture is internally consistent. It says nothing about whether the cases are representative of real inputs. Those are different questions with different tools.

Three boundaries to hold:

  • Token and substring detection misses semantic duplicates and paraphrased leakage.
  • Contamination inside a model's pretraining data is a different problem with different tooling. Auditing your own JSONL files cannot resolve it.
  • A clean audit is a necessary condition for trusting a score, not a sufficient one.

The honest output of a clean audit is a qualified claim: on this fixture, after this audit, the improvement holds for these cases.

The Decision Rule

Before you quote an eval number, run the audit and attach the ledger. The number travels with its evidence, or it does not travel.

Your next move is not a new fixture. Audit the one you already have. Run the checker, then mutate one record and run it again, so the procedure becomes yours rather than this article's. When you can predict what a mutation will produce before the checker confirms it, you have stopped memorizing a fixture and started owning a method.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

A duplicated eval case contributes to a claim that accuracy improved broadly across edge cases. Which ledger response best follows the article's decision process?
Question 1 of 2Comparison Reasoning

Focus: Choose a bounded ledger action that connects a finding to the specific claim it compromises.

A paraphrased prompt shares meaning with a development-history example but few tokens, and the audit reports no match. What conclusion is justified?
Question 2 of 2Misconception Check

Focus: Recognize that a clean token- or substring-based audit cannot rule out semantic duplicates or paraphrased leakage.

References

  1. Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs - ACL Anthologyaclanthology.org
  2. Judging by the Cover: Cleaning LLM Truthfulness Benchmarks to Avoid Surface-Level Feature Leakagearxiv.org
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.