Skip to content
intermediate

Practice Auditing Coverage in an LLM Evaluation Suite

A suite that passes every case can still fail in production. Not because the model is weak, but because the cases never mapped to the task space you…

Published 2026-10-03Updated 2026-10-0411 min read
Captivating view of sand dunes and sparse vegetation under a subtle sunset sky.
Captivating view of sand dunes and sparse vegetation under a subtle sunset sky. Photo by Jacob Moore on Pexels.

A suite that passes every case can still fail in production. Not because the model is weak, but because the cases never mapped to the task space you actually ship into. Coverage is a counting problem you can run, not a feeling you argue about.

This exercise gives you one fixture, one script, one report, and one bounded set of additions. You will build a coverage audit that reconciles case totals, exposes uncovered strata, and flags overrepresented regions. Then you will propose a small change set where every addition names the gap it closes.

If you already know how to author a single evaluation case with expected criteria and failure labels, you have the prerequisite. This piece audits the shape of an existing suite, not how to write one case.

Why a Passing Suite Can Still Be Blind

A suite is a sample, not a census. Passing means the sampled region works. It says nothing about whether you sampled the right region.

Two failure shapes hide behind a green suite:

  • Uncovered strata. No case exercises a task type you will deploy into. The suite is silent there, and silence reads as success.
  • Overrepresented regions. Many near-duplicate cases inflate confidence. You feel thorough because the count is high, but the count is concentrated.

Consequential edge conditions matter more than raw case count. A rare input with a high cost of failure deserves deliberate representation, even if it appears once in a thousand requests.

The deliverable here is a reconciled report plus a bounded addition list, each addition tied to a named gap.

The Fixture: Task Inventory and Proposed Cases

The audit consumes two structures.

Task inventory. A list of task strata, each with an id, a short label, a deployment weight (expected share of real traffic), and the edge-condition tags that are consequential for that stratum.

Proposed cases. Each case carries an id, exactly one stratum, and zero or more edge tags drawn from a controlled vocabulary.

The fixture is deliberately seeded with problems:

  • One stratum has zero cases.
  • One stratum is heavily overrepresented.
  • One consequential edge condition is declared for a stratum but exercised by no case in that stratum.

Here is the shape of the data:

{
  "edge_vocabulary": ["missing_input", "ambiguous", "out_of_scope", "conflicting_context", "long_input"],
  "strata": [
    {"id": "s_billing", "label": "billing question", "weight": 0.30, "consequential_edges": ["missing_input", "conflicting_context"]},
    {"id": "s_refund", "label": "refund request", "weight": 0.25, "consequential_edges": ["ambiguous", "out_of_scope"]},
    {"id": "s_technical", "label": "technical troubleshooting", "weight": 0.25, "consequential_edges": ["long_input", "missing_input"]},
    {"id": "s_escalation", "label": "escalation to human", "weight": 0.20, "consequential_edges": ["conflicting_context", "out_of_scope"]}
  ],
  "cases": [
    {"id": "c01", "stratum": "s_billing", "edges": ["missing_input"]},
    {"id": "c02", "stratum": "s_billing", "edges": []},
    {"id": "c03", "stratum": "s_billing", "edges": ["conflicting_context"]},
    {"id": "c04", "stratum": "s_billing", "edges": []},
    {"id": "c05", "stratum": "s_billing", "edges": ["missing_input"]},
    {"id": "c06", "stratum": "s_refund", "edges": ["ambiguous"]},
    {"id": "c07", "stratum": "s_refund", "edges": []},
    {"id": "c08", "stratum": "s_technical", "edges": ["long_input"]}
  ]
}

Eight cases. Four strata. One stratum (s_escalation) has no cases at all. One stratum (s_billing) holds five of eight cases. The out_of_scope condition is declared consequential for s_refund and s_escalation, and no case exercises it in either.

State the reconciliation invariant now: every case maps to exactly one stratum, and the report's case total must equal the fixture's case count. If those two numbers disagree, the report is invalid, not merely incomplete.

Knowledge check

Check your understanding

Answer this question before you continue.

Which condition must hold for the fixture's case-count reconciliation to be valid?
Single Choice

Focus: State the reconciliation invariant that makes the fixture's coverage report trustworthy.

Set Up and Load the Audit

Dependencies: Python 3.10 or newer, standard library only. No API keys, no network, no external services.

Load the fixture and validate before any analysis runs. An audit that quietly drops unmapped cases produces a confident, wrong report.

import json
from collections import Counter

def load_fixture(path):
    with open(path) as f:
        data = json.load(f)

    vocab = set(data["edge_vocabulary"])
    strata = {s["id"]: s for s in data["strata"]}
    cases = data["cases"]

    for c in cases:
        if c["stratum"] not in strata:
            raise ValueError(f"case {c['id']} maps to unknown stratum {c['stratum']}")
        for tag in c["edges"]:
            if tag not in vocab:
                raise ValueError(f"case {c['id']} uses unknown edge tag {tag}")

    return strata, cases, vocab

Validation comes first because the alternative is silent corruption. If a case points at a stratum that does not exist, you want a loud failure, not a case that vanishes from the count.

A clean load prints something like:

Loaded 8 cases across 4 strata. 5 edge tags in vocabulary.

If that line disagrees with your fixture, stop. Fix the fixture before reading any report.

Knowledge check

Check your understanding

Answer this question before you continue.

A fixture case names a stratum ID absent from the inventory. What should the loader do?
Debugging

Focus: Diagnose how the fixture loader should handle a case that references an unknown stratum.

Map Cases to Strata and Edge Conditions

Now build the coverage map. Group cases by stratum, count them, and compute each stratum's share of the suite.

The unit of edge coverage is the stratum-edge pair, not the bare tag. A tag seen in one stratum does not make it covered in another. out_of_scope is consequential for both s_refund and s_escalation; a case tagged out_of_scope inside s_refund closes the refund pair and leaves the escalation pair wide open. Collapse the pair into a global tag set and you hide exactly the gap you are hunting.

def build_report(strata, cases, vocab):
    by_stratum = Counter(c["stratum"] for c in cases)
    total = len(cases)

    rows = []
    for sid, s in strata.items():
        count = by_stratum.get(sid, 0)
        share = count / total if total else 0.0
        rows.append({
            "stratum": sid,
            "count": count,
            "share": share,
            "weight": s["weight"],
            "gap": share - s["weight"],
        })

    required = set()
    for sid, s in strata.items():
        for tag in s["consequential_edges"]:
            required.add((sid, tag))

    exercised = set()
    for c in cases:
        for tag in c["edges"]:
            exercised.add((c["stratum"], tag))

    missing_pairs = required - exercised
    return rows, total, missing_pairs

Two design choices matter here.

First, edge coverage is counted per stratum-edge pair. A case can carry several tags, and the same tag can be consequential in several strata. Building the required set from (stratum, tag) pairs and the exercised set from each case's own stratum and tags keeps those relationships intact.

Second, the counting is deliberately simple and inspectable. Every number traces back to case ids. A coverage number you cannot trace is not evidence; it is a vibe with a decimal point.

Knowledge check

Check your understanding

Answer this question before you continue.

A new `s_refund` case is tagged `out_of_scope`. Which coverage conclusion follows?
Scenario Interpretation

Focus: Track consequential edge coverage separately for each stratum-edge pair.

Run the Report and Read the Output

A four-row matrix shows billing with five cases and a crowded marker, refund with two cases and a missing out-of-scope edge, technical troubleshooting with one case, and escalation with zero cases plus missing conflicting-context and out-of-scope edges.
Compare case counts by stratum, then check consequential edges within that same stratum; a tag covered elsewhere does not close the gap.

The report has three blocks. Print them, then read them. The reconciliation line must compare two independently derived totals, so compute the mapped count from the per-stratum rows rather than reusing the fixture length.

def print_report(rows, fixture_total, missing_pairs):
    mapped_total = sum(r["count"] for r in rows)
    status = "PASS" if mapped_total == fixture_total else "FAIL"
    print(f"Reconciliation: {mapped_total} cases mapped, {fixture_total} in fixture -> {status}")
    if mapped_total != fixture_total:
        raise ValueError("report does not reconcile with the fixture")
    print()
    print(f"{'stratum':<14}{'cases':>6}{'share':>8}{'weight':>8}{'gap':>8}")
    for r in rows:
        print(f"{r['stratum']:<14}{r['count']:>6}{r['share']:>8.2f}{r['weight']:>8.2f}{r['gap']:>+8.2f}")
    print()
    uncovered = [r["stratum"] for r in rows if r["count"] == 0]
    print(f"Uncovered strata: {uncovered or 'none'}")
    print(f"Missing consequential stratum-edge pairs: {sorted(missing_pairs) or 'none'}")

A realistic excerpt:

Reconciliation: 8 cases mapped, 8 in fixture -> PASS

stratum        cases   share  weight     gap
s_billing          5    0.62    0.30   +0.32
s_refund           2    0.25    0.25   +0.00
s_technical        1    0.12    0.25   -0.13
s_escalation       0    0.00    0.20   -0.20

Uncovered strata: ['s_escalation']
Missing consequential stratum-edge pairs: [('s_escalation', 'conflicting_context'), ('s_escalation', 'out_of_scope'), ('s_refund', 'out_of_scope')]

Read it in this order. The reconciliation line first: if it fails, nothing below it is trustworthy. Then the uncovered strata. Then the missing pairs. Then the gap column.

One interpretation rule: a large positive gap is not automatically good. s_billing at +0.32 does not mean billing is well covered. It means the suite is padded with billing cases, and the padding is buying you confidence you did not earn.

Picture the map as two columns: inventory strata on the left, suite cases on the right. s_escalation is an empty row. s_billing is a crowded one. The picture makes the problem obvious before the arithmetic does.

Knowledge check

Check your understanding

Answer this question before you continue.

The report shows `s_billing` with a gap of `+0.32`. What is the appropriate interpretation?
Output Prediction

Focus: Interpret a positive stratum gap as a comparison between suite share and deployment weight.

Common Mistakes and Debugging Signals

These are the errors that make a coverage audit lie to you.

  • Treating case count as coverage. A suite of 200 cases can still leave a stratum empty. Count per stratum, not per suite.
  • Collapsing edge tags into a global set. A tag exercised in one stratum is not covered in another. Track (stratum, tag) pairs, or you will report a gap as closed when it is still open.
  • Letting unmapped cases disappear. If the mapped total is less than the fixture total, the report is invalid. Do not round it away.
  • Confusing edge tags with edge cases. Tagging a case ambiguous does not mean the ambiguity is consequential for that stratum. Only the inventory decides what is consequential.
  • Over-trusting deployment weights. If the weights are guesses, the gap column is a guess with arithmetic applied to it. Label them as assumptions.
  • Chasing the loudest number. A stratum with zero cases and a high deployment weight is the loudest line in the report. Fix that before tuning anything else.

The debugging signal to watch: any stratum where count == 0 and weight > 0.10. That combination is a blind spot with real traffic behind it.

Propose Bounded Additions Tied to Named Gaps

Turn the report into a small, defensible change set. Two rules:

  1. Every proposed addition names the gap it closes, either a stratum id or a specific stratum-edge pair.
  2. Bound the set. Add a fixed small number of cases, prioritized by deployment weight times cost of failure, not by how easy the case is to write.

For the fixture above, a bounded set of three additions:

  • Add one case to s_escalation tagged conflicting_context. Closes the uncovered stratum with the second-highest weight and the (s_escalation, conflicting_context) pair.
  • Add one case to s_technical tagged missing_input. Closes the under-covered stratum and exercises a consequential edge.
  • Add one case to s_refund tagged out_of_scope. Closes the (s_refund, out_of_scope) pair.

Notice what is not here: no new billing cases. The overrepresented region needs removal or consolidation, not addition. State that explicitly in your change set, because "add more cases" is the reflex that created the imbalance.

Also notice what the third addition does not close. (s_escalation, out_of_scope) stays open, because a case in s_refund cannot exercise an escalation. If you want that pair closed, you need a fourth addition inside s_escalation—or you accept the gap and say so. Either is fine. Claiming the refund case closed it is not.

After applying the additions, rerun the audit. The expected post-addition report shows s_escalation at one case, s_technical at two, and the missing-pair list down to [('s_escalation', 'out_of_scope')]. The reconciliation line still reads PASS.

Modify the Experiment: Change the Weights, Watch the Gaps Move

Coverage is relative to a deployment assumption, not an absolute property. Prove it to yourself.

Change the deployment weights to a different plausible population. Say you learn that escalations are rarer than you thought and technical troubleshooting is more common:

strata["s_escalation"]["weight"] = 0.05
strata["s_technical"]["weight"] = 0.40

Rerun the same script without touching the cases. Watch s_technical flip from under-covered to badly under-covered, and watch the escalation addition stop being justified by weight alone. The cases did not change. The verdict did.

Second variation: add one tag to the controlled vocabulary, say non_english, and declare it consequential for s_billing. Rerun. Coverage that looked complete a minute ago is now incomplete. This is why every coverage claim needs the population assumption attached to it.

What This Audit Does Not Prove

Coverage is a structural check on the suite. It says nothing about whether the system passes those cases. A suite can be perfectly balanced and still fail every one.

A synthetic inventory is a model of the task space. A wrong inventory produces a tidy, confident, wrong report. The audit cannot tell you the inventory is wrong; only production traffic can.

Small suites cannot establish representativeness. They can show that named regions are present or absent. That is a weaker claim than "this suite represents our users," and you should not upgrade it in the writeup.

Coverage counts also do not capture difficulty, ordering effects, or interaction between conditions. A case can be present and still be too easy to reveal anything.

Use the audit to find blind spots and justify additions. Use execution results to judge quality. They are different jobs.

Where to Go Next

Before trusting any suite result, run the coverage audit and read the uncovered-strata and missing-pair lines first. If those lines are not empty, the pass rate is answering a question you did not ask.

Take your bounded addition list, add the cases, and rerun the audit to confirm the invariant holds and the named gaps closed. Then move to executing the suite and labeling failures. Coverage tells you where you are looking. Execution tells you what you found.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

Which summary accurately describes the article's three proposed additions and their remaining gap?
Question 1 of 2Comparison Reasoning

Focus: Verify that a bounded addition plan closes named gaps without claiming a pair was covered in the wrong stratum.

After a small suite has balanced stratum counts and its report reconciles, which conclusion is justified?
Question 2 of 2Misconception Check

Focus: Distinguish structural coverage evidence from proof of suite representativeness or system quality.

References

  1. Evaluation best practices | OpenAI APIdevelopers.openai.com
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.