Skip to content
intermediate

Measure Citation Coverage and Support in RAG Answers

A citation marker is a promise. The sentence says this passage backs me up. Most RAG evaluations never check whether the promise holds — they check whether…

Published 2026-10-03Updated 2026-10-0411 min read
Detailed texture of coarse sand dunes, perfect for backgrounds and design elements.
Detailed texture of coarse sand dunes, perfect for backgrounds and design elements. Photo by David Iloba on Pexels.

A citation marker is a promise. The sentence says this passage backs me up. Most RAG evaluations never check whether the promise holds — they check whether a marker exists.

Here is the failure that hides behind that check. You open an answer, and every sentence carries a bracketed number. You follow the third one. The passage is about the same topic, mentions the same product, uses the same vocabulary — and never states what the sentence claims. The marker passed your test. The claim did not.

That gap is what this tutorial measures. We will define two claim-level summaries, classify a small fixture by hand, compute both numbers, and then draw a hard line between these metrics and the retrieval precision and recall you already track.

Why Citation Presence Is the Wrong Unit

The default check most teams run is binary: does a citation marker exist next to this sentence? It is cheap, it is automatable, and it passes on the most common failure mode in citation prompting — post-hoc rationalization. The model generates a claim from parametric memory, then scans the retrieved context for something that looks compatible and attaches it. The marker is real. The grounding is theater.

Grant the narrow case where presence-checking is fine. Short answers. Single-source questions. Internal demos where nobody audits the output and the cost of a wrong claim is a shrug. If your answers are three sentences long and drawn from one document, presence is a reasonable proxy and you should not build a second eval harness for it.

The moment answers get longer, sources multiply, or a downstream agent acts on the output, presence stops discriminating. Consider one sentence:

The retry limit defaults to five attempts.

The cited span reads: "The client supports configurable retry behavior for transient failures." Topically adjacent. Same subsystem. Same document. Does not entail the number five. Presence is satisfied. Support is not.

The reframe is this: the measurable unit is the claim, and the measurable relation is whether the cited span entails it. Not whether the document is relevant. Not whether the topic matches. Whether the span, read on its own, makes the claim true.

Common mistake: Checking document-level relevance and calling it citation evaluation. If your check only asks "is this document about the right thing," you are measuring retrieval quality a second time, not citation faithfulness.

Knowledge check

Check your understanding

Answer this question before you continue.

An answer says, “The retry limit defaults to five attempts,” and cites a span that only says the client supports configurable retry behavior. What does a citation-presence check establish?
Misconception Check

Focus: Distinguish the existence of a citation marker from whether its cited span entails the claim.

Notation and the Two Denominators

Fix a fixture. An answer decomposes into atomic claims:

C={c1,c2,…,cn}C = \{c_1, c_2, \dots, c_n\}

Each claim cic_i carries zero or more cited spans — the specific passages the answer points at, not the whole retrieved document. Each claim gets exactly one label:

LabelCondition
SUPPORTEDAt least one cited span entails the claim
UNSUPPORTEDCitations exist, but none entails the claim
UNCITEDNo citation attached

The support label is all-or-nothing. A span must entail the claim as written to count as SUPPORTED. If the span supports a weaker version — a narrower error class, a default value instead of a configurable range — the claim is UNSUPPORTED. This rule only stays honest if claims are atomic: one checkable assertion each. A sentence that bundles two assertions has no clean slot in this scheme, so split it before you label. "Authentication failures are never retried and surface immediately" is two claims, and they may earn different labels.

Now the two summaries. They differ only in their denominator, and that difference is the whole point.

Citation coverage is the fraction of claims carrying at least one citation:

coverage=∣{ci:ci is cited}∣n\text{coverage} = \frac{|\{c_i : c_i \text{ is cited}\}|}{n}

This is approximately citation recall. It answers: how much of the answer is even attempting to show its work?

Citation support is the fraction of cited claims whose citation actually entails the claim:

support=∣{ci:ci is SUPPORTED}∣∣{ci:ci is cited}∣\text{support} = \frac{|\{c_i : c_i \text{ is SUPPORTED}\}|}{|\{c_i : c_i \text{ is cited}\}|}

This is the precision-flavored side. It answers: when the answer does cite, does the citation hold up?

Two assumptions carry the whole model. First, claims are atomic — one checkable assertion each. Second, claims are independently checkable — labeling one does not depend on labeling another. Break the first assumption and the denominator silently changes: a claim bundling two assertions can be half-supported, and your label set has no honest slot for it. Break the second and your counts drift as you reorder your review.

Note what both denominators count: claims. Not documents, not passages, not tokens. That is the first place these metrics part ways with retrieval metrics.

Knowledge check

Check your understanding

Answer this question before you continue.

A claim says, “Authentication failures are never retried.” Its cited span says permanent errors, such as authentication failures, are not retried. Which label follows the article’s rule?
Single Choice

Focus: Apply the all-or-nothing support label when a cited span supports only a weaker version of a claim.

Worked Example: Classifying Five Claims

Here is a fixture small enough to hold in your head. The question: "How does the sync client handle failures?" Three retrieved passages, abbreviated:

  • P1: "The sync client retries transient network failures up to five times with exponential backoff."
  • P2: "Permanent errors, such as authentication failures, are not retried and surface immediately to the caller."
  • P3: "Sync runs are logged to the local audit file for later inspection."

The generated answer, decomposed into five claims:

  1. The sync client retries transient failures up to five times. [P1]
  2. Retries use exponential backoff. [P1]
  3. Authentication failures are never retried. [P2]
  4. Sync runs are logged locally. [P3]
  5. The default retry limit can be raised through configuration. [P1]

Now classify each one. Read the claim. Read the cited span. Decide.

Claim 1 — SUPPORTED. P1 states the retry count directly. The span entails the claim with no inference required.

Claim 2 — SUPPORTED. P1 names exponential backoff as the retry strategy. Direct entailment.

Claim 3 — UNSUPPORTED. P2 says authentication failures are not retried. The claim says never. P2 describes one class of permanent error; it does not establish that no authentication failure is ever retried under any condition. The span supports a weaker claim than the one written. This is the partial-support case, and it lands as UNSUPPORTED because the label asks whether the span entails the claim as written.

Claim 4 — SUPPORTED. P3 states it plainly.

Claim 5 — UNSUPPORTED. P1 gives the default of five. It says nothing about configuration. The claim may well be true — the client probably does allow it — but the cited span does not entail it. This is the case that separates correctness from support: a true claim with a citation that does not carry it.

The label table is the artifact you keep:

#ClaimCited spanLabelJustification
1Retries transient failures up to five timesP1SUPPORTEDP1 states the count
2Uses exponential backoffP1SUPPORTEDP1 names the strategy
3Authentication failures are never retriedP2UNSUPPORTEDP2 covers one error class, not "never"
4Sync runs are logged locallyP3SUPPORTEDP3 states it
5Retry limit is configurableP1UNSUPPORTEDP1 gives the default only

Five claims. All five carry citations. Three are supported.

Calculating Coverage and Support Step by Step

Two rows compare the five claims: coverage marks all five as cited and shows 5 of 5, while support marks three as supported and shows 3 of 5 cited claims.
Coverage counts claims with citations; support counts cited claims whose spans entail them.

Count first, then divide.

Cited claims: every claim carries a citation — claims 1, 2, 3, 4, and 5. Cited count = 5.

Coverage:

coverage=55=1.0=100%\text{coverage} = \frac{5}{5} = 1.0 = 100\%

Every claim attempts to show its work.

Supported claims among cited claims: claims 1, 2, and 4. Supported count = 3.

Support:

support=35=0.6=60%\text{support} = \frac{3}{5} = 0.6 = 60\%

Read the pair together. Coverage of 100% with support of 60% says: this answer cites liberally, and two of its five citations do not hold up. If you had reported only coverage, you would have shipped a perfect score on an answer where 40% of the citations are decorative.

The two numbers move independently, and that is why reporting one is misleading:

PatternWhat it means
High coverage, low supportThe model cites everything and verifies nothing
Low coverage, high supportThe model is honest but under-cites — claims go unmarked
High coverage, high supportCitation hygiene is good — still not proof of truth
Low coverage, low supportThe answer is largely ungrounded and unmarked

One aggregation caveat before you wire this into a dashboard. Averaging per-answer scores weights every answer equally. A one-claim answer and a twenty-claim answer contribute the same amount to the mean. If you want the summary to reflect the corpus rather than the answer count, pool the claim counts first and divide once.

Knowledge check

Check your understanding

Answer this question before you continue.

An answer has six atomic claims. All six have citations, and four cited spans entail their claims. What are its citation coverage and citation support, respectively?
Output Prediction

Focus: Calculate citation coverage and citation support from cited and supported claim counts.

Why These Are Not Retrieval Precision or Recall

You already know how to compute precision and recall over a ranked document list. Those metrics score retrieval: given a query, how well did the ranked list match relevance judgments? Citation metrics score something else entirely — the generated answer's claims against the spans the model chose to cite.

Retrieval metricsCitation metrics
What is scoredRanked document listClaims in the generated answer
Denominator countsDocumentsClaims
Failure revealedMissed or misranked evidenceCited spans that do not entail their claims
Failure missedWhether the model used what it retrievedWhether retrieval found the right evidence at all

The two can diverge in both directions. A system with perfect retrieval recall can score poorly on support: the right passage sat in context, and the model cited the wrong span or ignored it. A system with mediocre retrieval precision can score well on support: the model cited only the passages it actually used, and every one of them held up.

If you need the mechanics of precision, recall, reciprocal rank, and NDCG, that derivation lives in the retrieval-metrics material. Here, the boundary is the point: retrieval metrics tell you whether the evidence was available. Citation metrics tell you whether the answer used it honestly.

Knowledge check

Check your understanding

Answer this question before you continue.

A system retrieves the needed evidence, but its answer cites spans that do not entail several claims. Which interpretation matches the article?
Comparison Reasoning

Focus: Explain why successful retrieval does not guarantee high citation support.

What These Numbers Do Not Prove

Support is not correctness. A claim can be perfectly entailed by its cited span and still be false — because the source is wrong, outdated, or itself unverified. You measured the link between claim and citation. You did not measure the link between citation and reality.

Support is not faithfulness either. A post-rationalized citation can be technically entailing while the model never used the passage during generation. At the output level, a faithful citation and a lucky one look identical. Distinguishing them requires inspecting model internals, not the text.

Labeling is judgment, and the hard cases are exactly where you need it most: partial support, numerical claims, and long spans that drift away from the claim. An automated entailment labeler shares failure modes with the generator — both are language models reading the same kind of text. The cheapest guardrail is a human spot-check on borderline claims, not a bigger model.

Warning: Treat coverage and support as a signal about citation hygiene, not as a certificate of truth. A 100% support score means every citation entails its claim. It does not mean the answer is right.

Building the Fixture Into Your Own Pipeline

Start with ten to twenty real answers from your own system, not a synthetic benchmark. Hand-label the claims once. That pass calibrates your judgment before you automate anything, and it usually reveals that your own definition of "supported" was looser than you thought.

Keep the label table as the durable artifact: claim text, cited span, label, one-line justification. It is what you re-run when the prompt, the chunking, or the model changes. Compare the tables, not just the summary numbers — a coverage drop from 90% to 85% hides which claims lost their citations.

Track coverage and support as a pair, and watch the gap between them. A widening gap means the model is citing more and supporting less, which is the signature of a prompt that rewards markers over grounding.

Re-run the fixture after every change to chunking, citation prompting, or the generation model. Then, when the hand-labeling gets tedious, that tedium is your signal: you now know what a trustworthy label looks like, and you are ready to test whether an automated entailment check reproduces it on the borderline cases. Do not automate before you can grade the automation.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

A claim is entailed by its cited span, producing a supported label. What does that result establish by itself?
Question 1 of 2Misconception Check

Focus: Distinguish claim-to-citation support from the factual correctness of the claim and its source.

A corpus summary should reflect claim counts rather than give a one-claim answer the same weight as a twenty-claim answer. Which approach does the article recommend?
Question 2 of 2Scenario Interpretation

Focus: Choose an aggregation method that weights claim-level results across a corpus rather than weighting each answer equally.

References

  1. Review metrics for RAG evaluations that use LLMs (console) - Amazon Bedrockdocs.aws.amazon.com
  2. Why Your RAG Citations Are Lying: Post-Hoc Rationalization in Source Attributiontianpan.co
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.