Practice Resolving Conflicting Sources in a RAG Answer
Retrieval worked. Three chunks came back, all relevant, all cited. Two of them disagree. Now what?

Key topics
Retrieval worked. Three chunks came back, all relevant, all cited. Two of them disagree. Now what?
This is the failure mode that survives every fix you made to the retrieval stage. Your pipeline can return the right documents and still produce a wrong answer, because "right document" and "consistent document set" are different properties. A grounded answer links each claim to a supporting span. But when the spans contradict each other, grounding alone tells you nothing about what to write.
Research on knowledge conflicts in RAG systems points to the same conclusion: models often struggle to resolve conflicts between sources appropriately, and explicitly reasoning about the conflict improves response quality more than silently picking a winner. That's the skill this exercise builds — by hand, on a small passage set, before you try to automate it.
You'll work through three disputed claims, record evidence in a claim ledger, and choose one of four verdicts for each. Then you'll score your answer against a rubric and re-run the exercise with one variable changed.
The Four Verdicts and What Each One Costs
Before touching the passages, fix the output vocabulary. Every disputed claim gets exactly one verdict:
| Verdict | When it applies | What the answer looks like |
|---|---|---|
| Support | Evidence converges | State the claim plainly, cite the spans |
| Qualify | Evidence converges only under a condition | State the claim with its condition attached |
| Clarify | The disagreement is real and user intent decides which branch matters | Ask a specific question instead of guessing |
| Abstain | The passages cannot settle the claim and no clarification would help | Say what is unresolved and what would resolve it |
Qualifying and abstaining feel like weaker answers. They are often the correct ones. A confident wrong resolution is the expensive failure — it ships a settled-sounding claim that your own retrieved evidence contradicts.
These verdicts are per claim, not per answer. One response can support one claim, qualify a second, and abstain on a third. That's not indecision. That's precision.
Knowledge check
Check your understanding
Answer this question before you continue.
Conflict Types You Will Meet in the Passage Set
Stop treating all disagreements as the same problem. The conflict type determines the default verdict, and the type is usually visible in metadata you already have.
- Freshness conflict. Same fact, different dates. A revised figure supersedes an earlier release. Default: support, using the newer value, acknowledging the older one.
- Scope conflict. Both statements are true but cover different populations, regions, time windows, or product versions. Default: qualify, stating the claim per scope.
- Authority conflict. A primary document versus a secondary summary that paraphrases it loosely. Default: support, preferring the primary — but check whether the summary actually contradicts or merely compresses.
- Genuine disagreement. Expert opinion or methodology differences where recency and rank are irrelevant. Default: clarify or abstain. Never resolve by rank.
- Missing context. A passage that is relevant but silent on the specific claim. This is not a contradiction. Do not treat silence as disagreement.
Note: These defaults are starting points, not rules. The whole point of the exercise is to justify each verdict from the passages rather than from a lookup table.
Knowledge check
Check your understanding
Answer this question before you continue.
Set Up the Claim Ledger
The ledger turns a vague judgment call into an inspectable process. One row per claim — not per passage:
| Claim | Supporting spans | Conflicting spans | Conflict type | Verdict | Justification |
|---|---|---|---|---|---|
| Product X costs $49/mo | P2 (Apr revision) | P1 (Jan release, $39) | Freshness | Support | April revision supersedes January figure for the same metric |
Why a ledger? Because it forces you to notice when you have zero supporting spans. That's the most common accidental hallucination path: a claim that feels true, cites something relevant, and is supported by nothing.
Keep the justification to one sentence. Long justifications hide the reasoning. If you can't compress it, you haven't decided yet.
One scale note: pairwise comparison grows quadratically with the number of retrieved chunks. For three to twenty chunks, comparing everything is fine. Beyond that, you cluster or pre-index known conflict pairs rather than comparing all combinations.
Knowledge check
Check your understanding
Answer this question before you continue.
Worked Passages: Three Claims, Three Different Verdicts
Here is the synthetic passage set. Five chunks, each tagged with a date, a scope note, and a source type.
- P1 — Product overview, January. "Product X is priced at $39/month for the standard tier."
- P2 — Pricing revision, April. "Effective April, Product X standard tier is $49/month."
- P3 — Support policy, undated, region tag: Global. "Support hours are 8am–6pm, Monday through Friday."
- P4 — Support policy, undated, region tag: EU. "Support hours are 9am–5pm, Monday through Friday."
- P5 — Opinion column. "Product X is overpriced relative to its feature set."
Claim 1: "Product X costs $49/month" — Freshness
Ledger row before verdict:
| Claim | Supporting | Conflicting | Type | Verdict | Justification |
|---|---|---|---|---|---|
| Product X costs $49/mo | P2 | P1 | Freshness | ? | ? |
Both passages state a price for the same tier. P2 is dated April and explicitly revises the earlier figure. The conflict is temporal, so recency is the right tiebreaker here.
Verdict: Support. State $49/month as current, and note that the January figure was $39 before the revision.
The trap: P1 may carry the higher similarity score, because "Product X is priced at $39/month" matches a pricing query just as well as the revision does. If you resolved by score alone, you'd cite the stale number. This is exactly why score cannot be the universal tiebreaker — it measures relevance, not currency.
Claim 2: "Support hours are 8am–6pm" — Scope
Ledger row before verdict:
| Claim | Supporting | Conflicting | Type | Verdict | Justification |
|---|---|---|---|---|---|
| Support hours are 8am–6pm | P3 | P4 | Scope | ? | ? |
P3 and P4 look contradictory until you read the region tags. P3 is tagged Global; P4 is tagged EU. Both statements are true for their own scope.
Verdict: Qualify. "Support hours are 8am–6pm globally, and 9am–5pm in the EU."
A scope difference is not a factual conflict. If you'd resolved this by picking the higher-ranked source, you'd have silently dropped a region's actual hours.
Claim 3: "Product X is overpriced" — Genuine Disagreement
Ledger row before verdict:
| Claim | Supporting | Conflicting | Type | Verdict | Justification |
|---|---|---|---|---|---|
| Product X is overpriced | P5 | none | Genuine disagreement | ? | ? |
P5 is an opinion. The pricing passages (P1, P2) are silent on value — they state numbers, not judgments. Silence is not contradiction, so there's no conflicting span here, only an unsupported assertion.
Verdict: Clarify or abstain. Clarify only when the user can supply a comparison basis the system can actually evaluate — for example, "overpriced compared to which alternative?" If the user names a competitor or a budget, you can retrieve evidence for that comparison. If no clarification would bring in evidence the corpus lacks, abstain: "The retrieved passages state the price but contain no evidence on value for money."
The choice between clarify and abstain hinges on one question: would the clarification bring in evidence that could change the answer? If yes, ask. If no, abstain and say what would resolve it.
Score Your Answer Against the Rubric
Now check your work. The rubric rewards explicit handling of disagreement and penalizes unsupported resolution.
Rewarded:
- Naming the conflict explicitly ("the January and April figures differ")
- Citing both sides, not just the winner
- Attaching conditions to qualified claims
- Stating what would resolve an abstention
Penalized:
- Presenting a disputed claim as settled
- Citing only the winning span
- Resolving by rank or recency when the conflict is not temporal
- Inventing a resolution the passages do not support
Three self-check questions:
- Did I name the disagreement, or did I quietly pick a side?
- Did I justify the verdict from the passages, or from a rule I applied without checking?
- Would a reader know what is still uncertain after reading my answer?
Common mistake: Treating a silent passage as a contradicting one. P1 and P2 don't contradict P5 — they simply don't address value. Silence and disagreement require different verdicts.
Knowledge check
Check your understanding
Answer this question before you continue.
Change One Variable and Re-Run
Fixed exercises teach the procedure. Experiments teach the boundary. Re-run the ledger three times with one change each.
Experiment A: Remove the date metadata from P1 and P2. Re-decide Claim 1. Without dates, you can't tell which figure is current. Your verdict should shift to clarify or abstain — which proves the metadata was doing the work, not your judgment about the numbers themselves.
Experiment B: Swap the authority tags on P3 and P4. Re-decide Claim 2. If your verdict flips, you were ranking sources instead of reading scope. The region tag is the deciding fact, and authority doesn't change it.
Experiment C: Add a sixth passage agreeing with P5. Re-decide Claim 3. One opinion plus one supporting opinion is still opinion — but now you have a genuine disagreement between value judgments rather than an unsupported assertion. The verdict may shift from abstain to clarify.
Record what changed and why. The delta between your two ledgers is the actual lesson.
Where This Connects to a Real System
The manual ledger is the production version of conflict handling, minus the automation. Recency weighting, source-priority fields, and explicit conflict-reasoning instructions in the generation step are all doing what you just did by hand. Research shows that explicitly informing a model about the potential conflict category improves response appropriateness — which means the reasoning you practiced here is the same reasoning you'd encode into a prompt or a pipeline stage.
My rule: adjudicate per claim, name the conflict type before choosing a verdict, and treat rank and recency as inputs rather than verdicts. Rank tells you what's relevant. Recency tells you what's current. Neither tells you what's true when sources genuinely disagree.
The next practical step is instrumentation. Log every retrieved set where two chunks conflict, along with the conflict type and the verdict your system produced. After a week of real traffic, you'll know how often your own corpus generates this situation and which document pairs conflict most. That log is the ledger you just built — running on its own.
Knowledge check
Final check
Finish the article by checking the ideas you just learned.
References
Want a more structured LLMOps path?
Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.
Large Language Models Starter Pack
A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.
- 227-page Illustrated PDF edition
- 12 guided LLM engineering chapters
- Visual concept diagrams
- Self-assessment quizzes
- Bonus deep-dive sections
- Prompt design, structured output, context windows & RAG pipelines
- Agents, tool calling, prompt injection, evaluation & application lifecycles
Coming soon


