Practice Auditing Coverage in an LLM Evaluation Suite
A suite that passes every case can still fail in production. Not because the model is weak, but because the cases never mapped to the task space you…

Key topics
A suite that passes every case can still fail in production. Not because the model is weak, but because the cases never mapped to the task space you actually ship into. Coverage is a counting problem you can run, not a feeling you argue about.
This exercise gives you one fixture, one script, one report, and one bounded set of additions. You will build a coverage audit that reconciles case totals, exposes uncovered strata, and flags overrepresented regions. Then you will propose a small change set where every addition names the gap it closes.
If you already know how to author a single evaluation case with expected criteria and failure labels, you have the prerequisite. This piece audits the shape of an existing suite, not how to write one case.
Why a Passing Suite Can Still Be Blind
A suite is a sample, not a census. Passing means the sampled region works. It says nothing about whether you sampled the right region.
Two failure shapes hide behind a green suite:
- Uncovered strata. No case exercises a task type you will deploy into. The suite is silent there, and silence reads as success.
- Overrepresented regions. Many near-duplicate cases inflate confidence. You feel thorough because the count is high, but the count is concentrated.
Consequential edge conditions matter more than raw case count. A rare input with a high cost of failure deserves deliberate representation, even if it appears once in a thousand requests.
The deliverable here is a reconciled report plus a bounded addition list, each addition tied to a named gap.
The Fixture: Task Inventory and Proposed Cases
The audit consumes two structures.
Task inventory. A list of task strata, each with an id, a short label, a deployment weight (expected share of real traffic), and the edge-condition tags that are consequential for that stratum.
Proposed cases. Each case carries an id, exactly one stratum, and zero or more edge tags drawn from a controlled vocabulary.
The fixture is deliberately seeded with problems:
- One stratum has zero cases.
- One stratum is heavily overrepresented.
- One consequential edge condition is declared for a stratum but exercised by no case in that stratum.
Here is the shape of the data:
{
"edge_vocabulary": ["missing_input", "ambiguous", "out_of_scope", "conflicting_context", "long_input"],
"strata": [
{"id": "s_billing", "label": "billing question", "weight": 0.30, "consequential_edges": ["missing_input", "conflicting_context"]},
{"id": "s_refund", "label": "refund request", "weight": 0.25, "consequential_edges": ["ambiguous", "out_of_scope"]},
{"id": "s_technical", "label": "technical troubleshooting", "weight": 0.25, "consequential_edges": ["long_input", "missing_input"]},
{"id": "s_escalation", "label": "escalation to human", "weight": 0.20, "consequential_edges": ["conflicting_context", "out_of_scope"]}
],
"cases": [
{"id": "c01", "stratum": "s_billing", "edges": ["missing_input"]},
{"id": "c02", "stratum": "s_billing", "edges": []},
{"id": "c03", "stratum": "s_billing", "edges": ["conflicting_context"]},
{"id": "c04", "stratum": "s_billing", "edges": []},
{"id": "c05", "stratum": "s_billing", "edges": ["missing_input"]},
{"id": "c06", "stratum": "s_refund", "edges": ["ambiguous"]},
{"id": "c07", "stratum": "s_refund", "edges": []},
{"id": "c08", "stratum": "s_technical", "edges": ["long_input"]}
]
}
Eight cases. Four strata. One stratum (s_escalation) has no cases at all. One stratum (s_billing) holds five of eight cases. The out_of_scope condition is declared consequential for s_refund and s_escalation, and no case exercises it in either.
State the reconciliation invariant now: every case maps to exactly one stratum, and the report's case total must equal the fixture's case count. If those two numbers disagree, the report is invalid, not merely incomplete.
Knowledge check
Check your understanding
Answer this question before you continue.
Set Up and Load the Audit
Dependencies: Python 3.10 or newer, standard library only. No API keys, no network, no external services.
Load the fixture and validate before any analysis runs. An audit that quietly drops unmapped cases produces a confident, wrong report.
import json
from collections import Counter
def load_fixture(path):
with open(path) as f:
data = json.load(f)
vocab = set(data["edge_vocabulary"])
strata = {s["id"]: s for s in data["strata"]}
cases = data["cases"]
for c in cases:
if c["stratum"] not in strata:
raise ValueError(f"case {c['id']} maps to unknown stratum {c['stratum']}")
for tag in c["edges"]:
if tag not in vocab:
raise ValueError(f"case {c['id']} uses unknown edge tag {tag}")
return strata, cases, vocab
Validation comes first because the alternative is silent corruption. If a case points at a stratum that does not exist, you want a loud failure, not a case that vanishes from the count.
A clean load prints something like:
Loaded 8 cases across 4 strata. 5 edge tags in vocabulary.
If that line disagrees with your fixture, stop. Fix the fixture before reading any report.
Knowledge check
Check your understanding
Answer this question before you continue.
Map Cases to Strata and Edge Conditions
Now build the coverage map. Group cases by stratum, count them, and compute each stratum's share of the suite.
The unit of edge coverage is the stratum-edge pair, not the bare tag. A tag seen in one stratum does not make it covered in another. out_of_scope is consequential for both s_refund and s_escalation; a case tagged out_of_scope inside s_refund closes the refund pair and leaves the escalation pair wide open. Collapse the pair into a global tag set and you hide exactly the gap you are hunting.
def build_report(strata, cases, vocab):
by_stratum = Counter(c["stratum"] for c in cases)
total = len(cases)
rows = []
for sid, s in strata.items():
count = by_stratum.get(sid, 0)
share = count / total if total else 0.0
rows.append({
"stratum": sid,
"count": count,
"share": share,
"weight": s["weight"],
"gap": share - s["weight"],
})
required = set()
for sid, s in strata.items():
for tag in s["consequential_edges"]:
required.add((sid, tag))
exercised = set()
for c in cases:
for tag in c["edges"]:
exercised.add((c["stratum"], tag))
missing_pairs = required - exercised
return rows, total, missing_pairs
Two design choices matter here.
First, edge coverage is counted per stratum-edge pair. A case can carry several tags, and the same tag can be consequential in several strata. Building the required set from (stratum, tag) pairs and the exercised set from each case's own stratum and tags keeps those relationships intact.
Second, the counting is deliberately simple and inspectable. Every number traces back to case ids. A coverage number you cannot trace is not evidence; it is a vibe with a decimal point.
Knowledge check
Check your understanding
Answer this question before you continue.
Run the Report and Read the Output
The report has three blocks. Print them, then read them. The reconciliation line must compare two independently derived totals, so compute the mapped count from the per-stratum rows rather than reusing the fixture length.
def print_report(rows, fixture_total, missing_pairs):
mapped_total = sum(r["count"] for r in rows)
status = "PASS" if mapped_total == fixture_total else "FAIL"
print(f"Reconciliation: {mapped_total} cases mapped, {fixture_total} in fixture -> {status}")
if mapped_total != fixture_total:
raise ValueError("report does not reconcile with the fixture")
print()
print(f"{'stratum':<14}{'cases':>6}{'share':>8}{'weight':>8}{'gap':>8}")
for r in rows:
print(f"{r['stratum']:<14}{r['count']:>6}{r['share']:>8.2f}{r['weight']:>8.2f}{r['gap']:>+8.2f}")
print()
uncovered = [r["stratum"] for r in rows if r["count"] == 0]
print(f"Uncovered strata: {uncovered or 'none'}")
print(f"Missing consequential stratum-edge pairs: {sorted(missing_pairs) or 'none'}")
A realistic excerpt:
Reconciliation: 8 cases mapped, 8 in fixture -> PASS
stratum cases share weight gap
s_billing 5 0.62 0.30 +0.32
s_refund 2 0.25 0.25 +0.00
s_technical 1 0.12 0.25 -0.13
s_escalation 0 0.00 0.20 -0.20
Uncovered strata: ['s_escalation']
Missing consequential stratum-edge pairs: [('s_escalation', 'conflicting_context'), ('s_escalation', 'out_of_scope'), ('s_refund', 'out_of_scope')]
Read it in this order. The reconciliation line first: if it fails, nothing below it is trustworthy. Then the uncovered strata. Then the missing pairs. Then the gap column.
One interpretation rule: a large positive gap is not automatically good. s_billing at +0.32 does not mean billing is well covered. It means the suite is padded with billing cases, and the padding is buying you confidence you did not earn.
Picture the map as two columns: inventory strata on the left, suite cases on the right. s_escalation is an empty row. s_billing is a crowded one. The picture makes the problem obvious before the arithmetic does.
Knowledge check
Check your understanding
Answer this question before you continue.
Common Mistakes and Debugging Signals
These are the errors that make a coverage audit lie to you.
- Treating case count as coverage. A suite of 200 cases can still leave a stratum empty. Count per stratum, not per suite.
- Collapsing edge tags into a global set. A tag exercised in one stratum is not covered in another. Track
(stratum, tag)pairs, or you will report a gap as closed when it is still open. - Letting unmapped cases disappear. If the mapped total is less than the fixture total, the report is invalid. Do not round it away.
- Confusing edge tags with edge cases. Tagging a case
ambiguousdoes not mean the ambiguity is consequential for that stratum. Only the inventory decides what is consequential. - Over-trusting deployment weights. If the weights are guesses, the gap column is a guess with arithmetic applied to it. Label them as assumptions.
- Chasing the loudest number. A stratum with zero cases and a high deployment weight is the loudest line in the report. Fix that before tuning anything else.
The debugging signal to watch: any stratum where count == 0 and weight > 0.10. That combination is a blind spot with real traffic behind it.
Propose Bounded Additions Tied to Named Gaps
Turn the report into a small, defensible change set. Two rules:
- Every proposed addition names the gap it closes, either a stratum id or a specific stratum-edge pair.
- Bound the set. Add a fixed small number of cases, prioritized by deployment weight times cost of failure, not by how easy the case is to write.
For the fixture above, a bounded set of three additions:
- Add one case to
s_escalationtaggedconflicting_context. Closes the uncovered stratum with the second-highest weight and the(s_escalation, conflicting_context)pair. - Add one case to
s_technicaltaggedmissing_input. Closes the under-covered stratum and exercises a consequential edge. - Add one case to
s_refundtaggedout_of_scope. Closes the(s_refund, out_of_scope)pair.
Notice what is not here: no new billing cases. The overrepresented region needs removal or consolidation, not addition. State that explicitly in your change set, because "add more cases" is the reflex that created the imbalance.
Also notice what the third addition does not close. (s_escalation, out_of_scope) stays open, because a case in s_refund cannot exercise an escalation. If you want that pair closed, you need a fourth addition inside s_escalation—or you accept the gap and say so. Either is fine. Claiming the refund case closed it is not.
After applying the additions, rerun the audit. The expected post-addition report shows s_escalation at one case, s_technical at two, and the missing-pair list down to [('s_escalation', 'out_of_scope')]. The reconciliation line still reads PASS.
Modify the Experiment: Change the Weights, Watch the Gaps Move
Coverage is relative to a deployment assumption, not an absolute property. Prove it to yourself.
Change the deployment weights to a different plausible population. Say you learn that escalations are rarer than you thought and technical troubleshooting is more common:
strata["s_escalation"]["weight"] = 0.05
strata["s_technical"]["weight"] = 0.40
Rerun the same script without touching the cases. Watch s_technical flip from under-covered to badly under-covered, and watch the escalation addition stop being justified by weight alone. The cases did not change. The verdict did.
Second variation: add one tag to the controlled vocabulary, say non_english, and declare it consequential for s_billing. Rerun. Coverage that looked complete a minute ago is now incomplete. This is why every coverage claim needs the population assumption attached to it.
What This Audit Does Not Prove
Coverage is a structural check on the suite. It says nothing about whether the system passes those cases. A suite can be perfectly balanced and still fail every one.
A synthetic inventory is a model of the task space. A wrong inventory produces a tidy, confident, wrong report. The audit cannot tell you the inventory is wrong; only production traffic can.
Small suites cannot establish representativeness. They can show that named regions are present or absent. That is a weaker claim than "this suite represents our users," and you should not upgrade it in the writeup.
Coverage counts also do not capture difficulty, ordering effects, or interaction between conditions. A case can be present and still be too easy to reveal anything.
Use the audit to find blind spots and justify additions. Use execution results to judge quality. They are different jobs.
Where to Go Next
Before trusting any suite result, run the coverage audit and read the uncovered-strata and missing-pair lines first. If those lines are not empty, the pass rate is answering a question you did not ask.
Take your bounded addition list, add the cases, and rerun the audit to confirm the invariant holds and the named gaps closed. Then move to executing the suite and labeling failures. Coverage tells you where you are looking. Execution tells you what you found.
Knowledge check
Final check
Finish the article by checking the ideas you just learned.
References
Want a more structured LLMOps path?
Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.
Large Language Models Starter Pack
A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.
- 227-page Illustrated PDF edition
- 12 guided LLM engineering chapters
- Visual concept diagrams
- Self-assessment quizzes
- Bonus deep-dive sections
- Prompt design, structured output, context windows & RAG pipelines
- Agents, tool calling, prompt injection, evaluation & application lifecycles
Coming soon


