Practice Auditing Evidence for RAG Answer Claims
A citation is a pointer, not a proof. This drill teaches you to stop trusting the pointer and start reading the span.

Key topics
A citation is a pointer, not a proof. This drill teaches you to stop trusting the pointer and start reading the span.
Most RAG answers fail quietly. Retrieval worked, the citations render, the prose reads clean — and one claim in the middle is simply wrong. You will not catch it by skimming. You catch it by decomposing the answer into claims, reading each cited span as a skeptic, and labeling what you find.
This is a claim-level audit drill. It assumes you already know how citations link claims to evidence and how faithfulness differs from retrieval quality. If those boundaries are fuzzy, fix that first — the drill depends on them.
Why a Citation Is Not Evidence
The default beginner model is simple: if a claim has a citation, it is grounded. That model survives until the first answer that cites a real, relevant, on-topic passage that does not actually say what the answer claims.
Relevance is not sufficiency. A passage can share entities, vocabulary, and subject matter with a claim and still fail to justify it. The citation points somewhere real. The claim still floats free.
Three labels replace the binary "hallucinated / not hallucinated" check, because that binary collapses three different problems into one:
- Supported — the cited span asserts the claim, directly or through a clearly required inference.
- Contradicted — the cited span asserts the opposite, or asserts something incompatible with the claim.
- Unsupported — the cited span is silent on the claim, or the claim requires a fact that no cited span provides.
Those three labels map to three failure shapes. Missing evidence means a required fact was never retrieved. Partial evidence means one hop of a multi-hop claim is absent. Contradicting evidence means the passage says the opposite of the answer.
Two misleading-citation patterns are worth naming before you start, because they are the ones that fool careful readers:
- Fabricated citation — the cited source does not contain the asserted fact at all.
- Grounded hallucination — the source was retrieved, is genuinely relevant, and the answer still misreads it.
Note: This is claim-level auditing, not answer-level scoring. One bad claim does not invalidate a good answer, and one good citation does not rescue a bad claim. Keep the unit of judgment small.
Knowledge check
Check your understanding
Answer this question before you continue.
Set Up the Audit: One Answer, Its Claims, Its Passages
You need three things on the table: one question, one generated answer with numbered claims, and the retrieved passages with visible spans. Keep the set small enough to hold in working memory — four to six claims and three to four passages is enough to expose every label.
Here is the working scenario.
Question: What is the refund window for annual plans, and does it differ for enterprise contracts?
Answer:
- Annual plans carry a 30-day refund window. [P1]
- Monthly plans carry a 14-day refund window. [P2]
- Enterprise contracts follow the same 30-day window as annual plans. [P1]
- Refunds are issued to the original payment method. [P3]
- The refund window starts on the billing date. [P1]
- Enterprise contracts can be refunded at any time before renewal. [P4]
Passages:
- P1: "Annual subscriptions may be cancelled for a full refund within 30 days of the initial billing date. The window is measured from the date the first invoice is issued."
- P2: "Monthly subscriptions are billed on a recurring basis. Cancellation stops future billing; amounts already charged are not refunded."
- P3: "Approved refunds are returned to the payment instrument used for the original purchase."
- P4: "Enterprise agreements are governed by the signed order form. Refund terms, if any, are specified per contract."
Now run four mechanical steps.
Step 1 — Decompose. Split the answer into atomic, independently checkable claims. A compound sentence hiding two claims behind one citation is the most common audit failure. "Annual plans get a 30-day window and enterprise gets the same" is two claims, not one.
Step 2 — Map. For each claim, list the passage it cites and the exact span you will read. A citation to a whole document is not auditable. A citation to a span is.
Step 3 — Read the span as a skeptic. Ask what the passage actually asserts, not what you expect it to assert. Watch for entity swaps, number changes, date shifts, and yes/no flips. These are the four ways a claim drifts from its evidence while still looking grounded.
Step 4 — Label and record. One label per claim, plus the span that drove the decision. If you cannot point to a span, the label is unsupported by default.
The artifact you produce is a claim-to-evidence table:
| Claim | Cited span | Verdict | Reason |
|---|---|---|---|
| 1 | P1, sentence 1 | Supported | Span states 30 days from billing date directly |
| 2 | P2, sentence 2 | Contradicted | Span says charged amounts are not refunded |
| 3 | P1, sentence 1 | Unsupported | P1 covers annual plans only; no enterprise terms |
| 4 | P3, sentence 1 | Supported | Span states original payment instrument |
| 5 | P1, sentence 2 | Supported | Span measures from invoice date |
| 6 | P4, sentence 2 | Unsupported | Span defers to order form; no blanket right stated |
That table is the whole drill. Everything below is about earning each row honestly.
Knowledge check
Check your understanding
Answer this question before you continue.
Worked Example: Labeling Six Claims
Walk each claim and notice how the label is earned — or not.
Claim 1 — Supported. P1 says annual subscriptions may be cancelled for a full refund within 30 days of the initial billing date. The claim says annual plans carry a 30-day refund window. Direct assertion, same entity, same number, same window. What earns the label is directness, not topical overlap. If P1 had merely discussed annual billing without stating a refund window, claim 1 would be unsupported.
Claim 2 — Contradicted. P2 says amounts already charged are not refunded. The claim says monthly plans carry a 14-day refund window. The span does not just fail to support the claim — it asserts the opposite. This is more dangerous than a missing citation, because it looks maximally grounded: the citation is real, the passage is on-topic, and a skimming reader will accept it.
Claim 3 — Unsupported. P1 is about annual plans. It says nothing about enterprise contracts. The claim asserts that enterprise follows the same window. The passage is topically adjacent — same product family, same refund topic — but never asserts the specific relation the claim needs. This is the case beginners most often mislabel as supported, because the passage feels close enough.
Claim 4 — Supported. P3 states that approved refunds are returned to the payment instrument used for the original purchase. The claim says refunds go to the original payment method. Same assertion, different wording. Supported.
Claim 5 — Supported. P1's second sentence measures the window from the date the first invoice is issued. The claim says the window starts on the billing date. If your system treats invoice date and billing date as the same event, this is a direct match. If it does not, flag the ambiguity and label unsupported rather than guessing.
Claim 6 — Unsupported. P4 says enterprise refund terms are specified per contract. The claim asserts a blanket right to refund at any time before renewal. The passage explicitly defers the question. This is a misleading citation in its purest form: the source exists, the source is relevant, and the source does not say this.
Common mistake: Treating a correct conclusion as a supported claim. If the answer is right but the cited span does not carry it, the label is still unsupported. You are auditing the evidence trail, not the answer's plausibility.
When a span is genuinely ambiguous — claim 5 is the honest example — say so and label it unsupported. Abstention is a valid audit outcome. Guessing converts your audit into noise.
Knowledge check
Check your understanding
Answer this question before you continue.
Common Mislabeling Mistakes
Five errors make audits unreliable. Each has a fix.
Topicality bias. You label supported because the passage is about the right subject. Fix: require the passage to assert the claim's specific relation, not just its topic. Claim 3 is the test case.
Citation-presence bias. You treat the existence of a citation as partial credit. Fix: the citation is a pointer; the span is the evidence. If you have not read the span, you have not audited the claim.
Charity reading. You fill in the missing hop yourself because the answer is probably true. Fix: audit the evidence trail, not your prior belief about the answer. Claim 6 is correct in spirit and unsupported in evidence.
Compound-claim smuggling. You assign one label to a sentence containing two claims, one supported and one not. Fix: split first, label second. Always.
Over-escalation. You flag every imperfect claim as a system failure. Fix: separate the labeling decision from the remediation decision. They are different jobs, and the next section handles the second one.
Warning: The asymmetry matters. A false "supported" ships a wrong claim to a user with a citation attached — the citation makes it more convincing, not less. A false "unsupported" costs you a second look. Bias your uncertainty toward the cheaper error.
From Label to Decision: Correct, Escalate, or Abstain
A label is not an action. Convert it with a simple rule.
- Supported → keep the claim.
- Contradicted → do not ship. Surface the conflict to the user or the reviewer.
- Unsupported → either retrieve more evidence or drop the claim.
For unsupported claims, the label alone does not tell you which remediation path applies. The cited spans do. There are two distinct causes:
- Retrieval problem — the evidence exists in the corpus but was not fetched. Claim 3 is a candidate: enterprise refund terms may live in a contract document that was never retrieved.
- Generation problem — the evidence was fetched and the model still overreached. Claim 6 is the candidate: P4 was retrieved, and the model asserted a right the passage explicitly defers.
The distinction determines what you fix. A retrieval problem is a chunking, embedding, or query problem. A generation problem is a prompt, decoding, or verification problem. Tuning the wrong one wastes the effort.
Bound the escalation. "Escalate" should mean something concrete: flag the specific claim, attach the span that drove the verdict, state the missing hop in one sentence, and hand it to a human reviewer or a re-retrieval step. Do not escalate the whole answer when one claim is at fault — that floods the reviewer and hides the signal.
Abstention is a legitimate output. An answer that says "the retrieved passages do not establish enterprise refund terms" is more useful than a confident claim with a decorative citation. Build the abstention path before you need it.
Knowledge check
Check your understanding
Answer this question before you continue.
Your Turn: Audit a Fresh Answer
Here is a second set with no labels shown. Same format, different content.
Question: Does the standard support rate limits per API key, and what happens when a key exceeds its limit?
Answer:
- The standard requires rate limits to be enforced per API key. [A]
- Exceeding a limit returns HTTP 429. [B]
- Limits reset on a rolling 60-second window. [A]
- Keys can be temporarily suspended after repeated violations. [C]
- The standard specifies a minimum limit of 100 requests per minute. [A]
Passages:
- A: "Implementations must enforce request limits on a per-key basis. The measurement window is implementation-defined; a rolling 60-second window is common but not required."
- B: "When a client exceeds its allotted rate, the service returns HTTP 429 (Too Many Requests)."
- C: "Repeated violations may result in temporary suspension of the offending key at the operator's discretion."
Produce a completed claim-to-evidence table: one label and one cited span per claim, plus a one-line remediation decision for any non-supported claim.
Your audit is sound when three things hold. Every label traces to a specific span. No compound claim is left unsplit. At least one non-supported claim is identified if one exists — and in this set, more than one does.
Then run the modification that teaches the real lesson: remove passage A from the set and re-audit. Watch which claims flip from supported to unsupported. That flip is the missing-hop signal. It shows you that sufficiency is a property of the evidence set, not of any single passage — a claim can be supported by the set and unsupported by the passage it happens to cite.
What to Carry Forward
Never accept a claim because it has a citation. Accept it because you found the span that asserts it. That single rule separates an audit from a vibe check.
The table is a reusable artifact, not a one-off exercise. Keep it. Once you can label claims reliably, you can tell a retrieval failure from a generation failure — and that is the difference between fixing the right thing and tuning the wrong one.
This week, take one real RAG answer you already have, decompose it, and audit it claim by claim. Keep the table. The skill compounds: the more claims you label, the faster you spot the citation that points somewhere real and proves nothing.
Knowledge check
Final check
Finish the article by checking the ideas you just learned.
References
Want a more structured LLMOps path?
Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.
Large Language Models Starter Pack
A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.
- 227-page Illustrated PDF edition
- 12 guided LLM engineering chapters
- Visual concept diagrams
- Self-assessment quizzes
- Bonus deep-dive sections
- Prompt design, structured output, context windows & RAG pipelines
- Agents, tool calling, prompt injection, evaluation & application lifecycles
Coming soon


