Measure Citation Coverage and Support in RAG Answers
A citation marker is a promise. The sentence says this passage backs me up. Most RAG evaluations never check whether the promise holds — they check whether…

Key topics
A citation marker is a promise. The sentence says this passage backs me up. Most RAG evaluations never check whether the promise holds — they check whether a marker exists.
Here is the failure that hides behind that check. You open an answer, and every sentence carries a bracketed number. You follow the third one. The passage is about the same topic, mentions the same product, uses the same vocabulary — and never states what the sentence claims. The marker passed your test. The claim did not.
That gap is what this tutorial measures. We will define two claim-level summaries, classify a small fixture by hand, compute both numbers, and then draw a hard line between these metrics and the retrieval precision and recall you already track.
Why Citation Presence Is the Wrong Unit
The default check most teams run is binary: does a citation marker exist next to this sentence? It is cheap, it is automatable, and it passes on the most common failure mode in citation prompting — post-hoc rationalization. The model generates a claim from parametric memory, then scans the retrieved context for something that looks compatible and attaches it. The marker is real. The grounding is theater.
Grant the narrow case where presence-checking is fine. Short answers. Single-source questions. Internal demos where nobody audits the output and the cost of a wrong claim is a shrug. If your answers are three sentences long and drawn from one document, presence is a reasonable proxy and you should not build a second eval harness for it.
The moment answers get longer, sources multiply, or a downstream agent acts on the output, presence stops discriminating. Consider one sentence:
The retry limit defaults to five attempts.
The cited span reads: "The client supports configurable retry behavior for transient failures." Topically adjacent. Same subsystem. Same document. Does not entail the number five. Presence is satisfied. Support is not.
The reframe is this: the measurable unit is the claim, and the measurable relation is whether the cited span entails it. Not whether the document is relevant. Not whether the topic matches. Whether the span, read on its own, makes the claim true.
Common mistake: Checking document-level relevance and calling it citation evaluation. If your check only asks "is this document about the right thing," you are measuring retrieval quality a second time, not citation faithfulness.
Knowledge check
Check your understanding
Answer this question before you continue.
Notation and the Two Denominators
Fix a fixture. An answer decomposes into atomic claims:
Each claim carries zero or more cited spans — the specific passages the answer points at, not the whole retrieved document. Each claim gets exactly one label:
| Label | Condition |
|---|---|
| SUPPORTED | At least one cited span entails the claim |
| UNSUPPORTED | Citations exist, but none entails the claim |
| UNCITED | No citation attached |
The support label is all-or-nothing. A span must entail the claim as written to count as SUPPORTED. If the span supports a weaker version — a narrower error class, a default value instead of a configurable range — the claim is UNSUPPORTED. This rule only stays honest if claims are atomic: one checkable assertion each. A sentence that bundles two assertions has no clean slot in this scheme, so split it before you label. "Authentication failures are never retried and surface immediately" is two claims, and they may earn different labels.
Now the two summaries. They differ only in their denominator, and that difference is the whole point.
Citation coverage is the fraction of claims carrying at least one citation:
This is approximately citation recall. It answers: how much of the answer is even attempting to show its work?
Citation support is the fraction of cited claims whose citation actually entails the claim:
This is the precision-flavored side. It answers: when the answer does cite, does the citation hold up?
Two assumptions carry the whole model. First, claims are atomic — one checkable assertion each. Second, claims are independently checkable — labeling one does not depend on labeling another. Break the first assumption and the denominator silently changes: a claim bundling two assertions can be half-supported, and your label set has no honest slot for it. Break the second and your counts drift as you reorder your review.
Note what both denominators count: claims. Not documents, not passages, not tokens. That is the first place these metrics part ways with retrieval metrics.
Knowledge check
Check your understanding
Answer this question before you continue.
Worked Example: Classifying Five Claims
Here is a fixture small enough to hold in your head. The question: "How does the sync client handle failures?" Three retrieved passages, abbreviated:
- P1: "The sync client retries transient network failures up to five times with exponential backoff."
- P2: "Permanent errors, such as authentication failures, are not retried and surface immediately to the caller."
- P3: "Sync runs are logged to the local audit file for later inspection."
The generated answer, decomposed into five claims:
- The sync client retries transient failures up to five times.
[P1] - Retries use exponential backoff.
[P1] - Authentication failures are never retried.
[P2] - Sync runs are logged locally.
[P3] - The default retry limit can be raised through configuration.
[P1]
Now classify each one. Read the claim. Read the cited span. Decide.
Claim 1 — SUPPORTED. P1 states the retry count directly. The span entails the claim with no inference required.
Claim 2 — SUPPORTED. P1 names exponential backoff as the retry strategy. Direct entailment.
Claim 3 — UNSUPPORTED. P2 says authentication failures are not retried. The claim says never. P2 describes one class of permanent error; it does not establish that no authentication failure is ever retried under any condition. The span supports a weaker claim than the one written. This is the partial-support case, and it lands as UNSUPPORTED because the label asks whether the span entails the claim as written.
Claim 4 — SUPPORTED. P3 states it plainly.
Claim 5 — UNSUPPORTED. P1 gives the default of five. It says nothing about configuration. The claim may well be true — the client probably does allow it — but the cited span does not entail it. This is the case that separates correctness from support: a true claim with a citation that does not carry it.
The label table is the artifact you keep:
| # | Claim | Cited span | Label | Justification |
|---|---|---|---|---|
| 1 | Retries transient failures up to five times | P1 | SUPPORTED | P1 states the count |
| 2 | Uses exponential backoff | P1 | SUPPORTED | P1 names the strategy |
| 3 | Authentication failures are never retried | P2 | UNSUPPORTED | P2 covers one error class, not "never" |
| 4 | Sync runs are logged locally | P3 | SUPPORTED | P3 states it |
| 5 | Retry limit is configurable | P1 | UNSUPPORTED | P1 gives the default only |
Five claims. All five carry citations. Three are supported.
Calculating Coverage and Support Step by Step
Count first, then divide.
Cited claims: every claim carries a citation — claims 1, 2, 3, 4, and 5. Cited count = 5.
Coverage:
Every claim attempts to show its work.
Supported claims among cited claims: claims 1, 2, and 4. Supported count = 3.
Support:
Read the pair together. Coverage of 100% with support of 60% says: this answer cites liberally, and two of its five citations do not hold up. If you had reported only coverage, you would have shipped a perfect score on an answer where 40% of the citations are decorative.
The two numbers move independently, and that is why reporting one is misleading:
| Pattern | What it means |
|---|---|
| High coverage, low support | The model cites everything and verifies nothing |
| Low coverage, high support | The model is honest but under-cites — claims go unmarked |
| High coverage, high support | Citation hygiene is good — still not proof of truth |
| Low coverage, low support | The answer is largely ungrounded and unmarked |
One aggregation caveat before you wire this into a dashboard. Averaging per-answer scores weights every answer equally. A one-claim answer and a twenty-claim answer contribute the same amount to the mean. If you want the summary to reflect the corpus rather than the answer count, pool the claim counts first and divide once.
Knowledge check
Check your understanding
Answer this question before you continue.
Why These Are Not Retrieval Precision or Recall
You already know how to compute precision and recall over a ranked document list. Those metrics score retrieval: given a query, how well did the ranked list match relevance judgments? Citation metrics score something else entirely — the generated answer's claims against the spans the model chose to cite.
| Retrieval metrics | Citation metrics | |
|---|---|---|
| What is scored | Ranked document list | Claims in the generated answer |
| Denominator counts | Documents | Claims |
| Failure revealed | Missed or misranked evidence | Cited spans that do not entail their claims |
| Failure missed | Whether the model used what it retrieved | Whether retrieval found the right evidence at all |
The two can diverge in both directions. A system with perfect retrieval recall can score poorly on support: the right passage sat in context, and the model cited the wrong span or ignored it. A system with mediocre retrieval precision can score well on support: the model cited only the passages it actually used, and every one of them held up.
If you need the mechanics of precision, recall, reciprocal rank, and NDCG, that derivation lives in the retrieval-metrics material. Here, the boundary is the point: retrieval metrics tell you whether the evidence was available. Citation metrics tell you whether the answer used it honestly.
Knowledge check
Check your understanding
Answer this question before you continue.
What These Numbers Do Not Prove
Support is not correctness. A claim can be perfectly entailed by its cited span and still be false — because the source is wrong, outdated, or itself unverified. You measured the link between claim and citation. You did not measure the link between citation and reality.
Support is not faithfulness either. A post-rationalized citation can be technically entailing while the model never used the passage during generation. At the output level, a faithful citation and a lucky one look identical. Distinguishing them requires inspecting model internals, not the text.
Labeling is judgment, and the hard cases are exactly where you need it most: partial support, numerical claims, and long spans that drift away from the claim. An automated entailment labeler shares failure modes with the generator — both are language models reading the same kind of text. The cheapest guardrail is a human spot-check on borderline claims, not a bigger model.
Warning: Treat coverage and support as a signal about citation hygiene, not as a certificate of truth. A 100% support score means every citation entails its claim. It does not mean the answer is right.
Building the Fixture Into Your Own Pipeline
Start with ten to twenty real answers from your own system, not a synthetic benchmark. Hand-label the claims once. That pass calibrates your judgment before you automate anything, and it usually reveals that your own definition of "supported" was looser than you thought.
Keep the label table as the durable artifact: claim text, cited span, label, one-line justification. It is what you re-run when the prompt, the chunking, or the model changes. Compare the tables, not just the summary numbers — a coverage drop from 90% to 85% hides which claims lost their citations.
Track coverage and support as a pair, and watch the gap between them. A widening gap means the model is citing more and supporting less, which is the signature of a prompt that rewards markers over grounding.
Re-run the fixture after every change to chunking, citation prompting, or the generation model. Then, when the hand-labeling gets tedious, that tedium is your signal: you now know what a trustworthy label looks like, and you are ready to test whether an automated entailment check reproduces it on the borderline cases. Do not automate before you can grade the automation.
Knowledge check
Final check
Finish the article by checking the ideas you just learned.
References
Want a more structured LLMOps path?
Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.
Large Language Models Starter Pack
A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.
- 227-page Illustrated PDF edition
- 12 guided LLM engineering chapters
- Visual concept diagrams
- Self-assessment quizzes
- Bonus deep-dive sections
- Prompt design, structured output, context windows & RAG pipelines
- Agents, tool calling, prompt injection, evaluation & application lifecycles
Coming soon


