Skip to content
intermediate

Test RAG Query Rewrites Against a Fixed Retrieval Task

You add a rewrite step. The demo answers the question. You feel clever. Then someone asks a slightly different question and the whole thing falls apart,…

Published 2026-10-03Updated 2026-10-0410 min read
Child diving underwater with sun rays filtering through, creating a magical aquatic scene.
Child diving underwater with sun rays filtering through, creating a magical aquatic scene. Photo by James Cheney on Pexels.

You add a rewrite step. The demo answers the question. You feel clever. Then someone asks a slightly different question and the whole thing falls apart, and you realize you never actually knew whether the rewrite was doing the work or the question was just easy.

Here is the fix: stop judging rewrites by how the final answer sounds. Put the rewrite on trial against a direct-query baseline, on one fixed corpus, with one dumb retriever, and let relevance labels decide the winner.

This is a drill, not a lecture. You will run a direct query, write your own bounded rewrite, run it through the same retriever, and fill in an evidence table. By the end you will have at least one case where your rewrite helped and one where it hurt.

If you already know when to rewrite, expand, or decompose, this exercise answers the harder question: did your rewrite earn its place?

Why a Fixed Retrieval Task Beats Vibes

A fixed corpus, retriever, and top-k feed two query paths: direct and rewritten. Their retrieved results are evaluated against the same relevance labels to produce a helped, hurt, or neutral verdict.
Hold the retrieval setup fixed so the query wording is the only changing variable.

The most common beginner mistake is changing three things at once. You swap the retriever, add documents, and rewrite the query, then credit the rewrite for the improvement. That is not an experiment. That is a coincidence with a narrative.

An experiment holds everything constant except the variable you care about. Here, the only thing that moves is the query text. Same corpus. Same retriever. Same top-k. If the score changes, the query caused it.

That is why the retriever in this drill is deliberately boring: a standard-library token-overlap scorer. No embeddings, no API keys, no network. It tokenizes, lowercases, counts shared terms, and ranks. It is deterministic and inspectable, which means any score change is attributable to the words you fed it.

Note: This measures retrieval, not answer quality. A rewrite can improve retrieval and still fail to improve the final answer, because the generator has its own failure modes. Keep those two questions separate or you will misattribute everything.

Relevance labels are the referee. Without them, "better results" is a feeling. With them, you can compute precision and recall by hand and force a verdict on every row.

Knowledge check

Check your understanding

Answer this question before you continue.

A team changes the query and swaps in a new retriever, while keeping the corpus and top-k fixed. Can it attribute a score change specifically to the rewrite?
Single Choice

Focus: Identify which experimental controls make a retrieval change attributable to the query rewrite.

Set Up the Corpus, Retriever, and Labels

The whole harness fits in one file. No dependencies beyond the standard library.

import re
from collections import Counter

CORPUS = {
    "d1": "The Standard Savings Account pays 2.5% annual interest and has no monthly fees. It is for domestic use only and does not include travel benefits.",
    "d2": "The Travel Rewards Card includes international usage, travel insurance, and airport lounge access. It charges an annual fee.",
    "d3": "The Personal Loan covers home improvement and auto purchases with terms from 12 to 84 months. Loans are separate from deposit accounts.",
    "d4": "The Premium Checking Account offers online and mobile banking, no minimum balance, and no maintenance fees.",
    "d5": "International wire transfers are available on Premium Checking and Travel Rewards products for a flat fee.",
}

def tokenize(text):
    return re.findall(r"[a-z0-9]+", text.lower())

def retrieve(query, k=3):
    q = Counter(tokenize(query))
    scored = []
    for doc_id, text in CORPUS.items():
        d = Counter(tokenize(text))
        score = sum(min(q[t], d[t]) for t in q)
        scored.append((doc_id, score))
    scored.sort(key=lambda x: (-x[1], x[0]))
    return scored[:k]

LABELS = {
    "Can I use it abroad?": {"d2", "d5"},
    "What are the loan terms?": {"d3"},
    "Does the savings account have travel benefits?": {"d1"},
}

The corpus has deliberate overlap: savings, checking, and travel products all mention fees, and two documents mention international usage. That overlap is what makes some questions genuinely ambiguous and others not.

The scorer uses an overlap count with a cap per term (min(q[t], d[t])), so repeating a word in the query does not inflate the score. Run the direct query first and look at the raw ranked output before you touch anything.

for question in LABELS:
    print(question)
    for doc_id, score in retrieve(question):
        print(f"  {doc_id}  score={score}")
    print()

Expected output for the first question looks roughly like this:

Can I use it abroad?
  d2  score=2
  d5  score=1
  d1  score=0

That ranked list plus a precision/recall number is your baseline. Every rewrite must clear that bar.

Knowledge check

Check your understanding

Answer this question before you continue.

In the article's overlap scorer, what happens if a query repeats a term that also appears in a document?
Misconception Check

Focus: Explain how the overlap scorer handles repeated query terms.

Write Bounded Rewrites and Decompositions

Bounded means three things: keep the same information need, change only the wording, and stay inside a stated length or term budget. An unbounded rewrite makes failures unattributable, because you cannot tell which of your ten new words broke it.

A rewrite resolves pronouns, swaps user vocabulary for corpus vocabulary, and drops conversational filler. A decomposition splits a multi-part question into sub-queries, retrieves for each, then merges the result sets. Merging changes what precision means, so define your merge rule before you run it — for example, "union of top-3 per sub-query, then re-rank by summed score."

Start with one rewrite per question. Two is the natural next step, not the first one.

Common mistake: Rewriting toward what you hope the corpus says instead of what the question asks. If you peek at the labels and then write a query containing the answer's keywords, you are measuring your own cheating, not the rewrite. Write the rewrite before you look at the ranked output.

Knowledge check

Check your understanding

Answer this question before you continue.

You are testing a rewrite for an ambiguous question. Which approach best follows the article's bounded-rewrite guidance?
Scenario Interpretation

Focus: Choose a rewrite strategy that preserves the information need while making retrieval changes interpretable.

Your Turn: Attempt Before You Look

Before reading any worked example, do the work yourself. For each question in LABELS, write one bounded rewrite on paper or in a scratch file. Do not run it yet. Do not look at the ranked output. Just write the query you believe will retrieve the relevant documents.

Then run each rewrite through the identical retriever and the same k. Change nothing else. Record the retrieved set, the relevant set, precision, recall, and a verdict.

QuestionQuery usedRetrievedRelevantPrecisionRecallVerdict
Can I use it abroad?directbaseline
Can I use it abroad?your rewrite
What are the loan terms?directbaseline
What are the loan terms?your rewrite
Does the savings account have travel benefits?directbaseline
Does the savings account have travel benefits?your rewrite

Precision is relevant-retrieved over retrieved. Recall is relevant-retrieved over relevant. Compute both by hand at k=3; the arithmetic is the point, because it forces you to look at what actually came back.

Force a verdict on every row: helped, hurt, or neutral. If you cannot decide, check three things before you blame the rewrite: are the relevance labels correct, did the retrieved set actually change, and does your verdict rule handle ties? Only tighten the rewrite if it changed the information need or exceeded your stated constraint.

Inspect the retrieved evidence, not just the score. A rewrite that raises recall by dragging in loosely related documents can be worse for the downstream answer, because the generator now has to ignore noise. The score is a proxy; the documents are the reality.

Worked Examples: One Helpful, One Harmful

Now compare your attempt against two worked cases. If your rewrite landed in the same verdict bucket, you are reading the mechanism correctly. If not, the diff will show you why.

A helpful case

For "Does the savings account have travel benefits?", the direct query retrieves:

Does the savings account have travel benefits?
  d1  score=3
  d2  score=1
  d5  score=0

The relevant document is d1, and it is already ranked first. The direct query is strong here because the question uses corpus vocabulary almost verbatim. A rewrite that swaps "travel benefits" for "travel insurance" would add a term that appears in d2, not d1, and could pull the wrong document up. This is a case where the direct query wins and a rewrite is likely to hurt.

Now try a question where the direct query is weak. For "Can I use it abroad?", the direct query retrieves d2 and d5 but the pronoun "it" carries no retrieval signal. A bounded rewrite that resolves the pronoun and uses corpus vocabulary — "international usage travel card" — keeps the same information need while giving the retriever real terms to match. The retrieved set stays d2, d5, d1, precision stays 0.67, and recall stays 1.00. The rewrite is neutral on this corpus, but it is neutral for the right reason: the direct query already found the relevant documents because "abroad" overlaps with "international."

Knowledge check

Check your understanding

Answer this question before you continue.

For “Can I use it abroad?”, the direct query and the rewrite “international usage travel card” both retrieve d2, d5, and d1 at k=3. The relevant documents are d2 and d5. What is the correct verdict based on these results?
Comparison Reasoning

Focus: Interpret a rewrite's effect using the retrieved evidence and relevance labels rather than assuming a wording change helped.

A harmful case

For "What are the loan terms?", the direct query retrieves d3 first. Now write a rewrite that adds synonyms the corpus never uses: "personal loan repayment period duration months". The corpus says "terms from 12 to 84 months," not "repayment period" or "duration." The token overlap for d3 drops because the new terms do not match, and d1 or d5 may climb into the top-3 on incidental matches. Precision falls. This is synonym drift: you replaced a match with a miss.

The lesson is not that rewriting is bad. The lesson is that a rewrite is a hypothesis about vocabulary, and the corpus is the judge.

Read the Harmful Cases

Failure is the most instructive part of this drill. Three patterns show up again and again.

Synonym drift. Your rewrite adds synonyms the corpus never uses. Token overlap drops, and the relevant document falls out of the top-k. If the corpus says "annual fee" and you write "yearly charge," you have replaced a match with a miss.

Over-decomposition. You split a question that was never ambiguous. Each sub-query retrieves a different slice, and the merged set has lower precision than the direct query. Decomposition earns its keep on multi-part questions, not on single-fact ones.

Verbosity. A longer, more elaborate rewrite can score worse than the plain question, because extra terms compete for weight. This is a documented failure mode in query-rewriting research: a rewrite that is technically correct but needlessly wordy can retrieve worse than the original.

The diagnostic habit that explains most of this: diff the token sets. Take the direct query's tokens, take the rewrite's tokens, and look at what you added and removed. Those additions and removals are the cause of nearly every score change you will see.

Record harmful cases with the same rigor as helpful ones. A rewrite that helps two questions and hurts three is a net loss, and you only know that if you logged both.

Extend the Experiment

Once the basic loop works, one modification compounds the value.

Add a second rewrite strategy per question and compare all three columns side by side. Now you are not asking "does rewriting help?" but "which rewrite helps, and on which questions?" That is the question that actually ships.

Then vary k. A gain that vanishes at k=10 was probably noise at k=3. Finally, swap in a different retriever — even a simple embedding-based one — and check whether your conclusions hold. Rewrites tuned to one retriever often do not transfer, because the rewrite was exploiting that retriever's quirks rather than the question's meaning.

Keep the harness. A fixed corpus, a dumb retriever, and a label file is a reusable test bench for every future rewrite idea you have.

What to Do Next

Do not ship a rewrite step you have not measured against a direct-query baseline on a fixed corpus with labels. That is the whole decision rule, and it fits in one sentence because the experiment is small enough to actually run.

The next move is mechanical: keep the harness, add a second rewrite strategy, and re-run whenever the corpus or the retriever changes. A rewrite that helped yesterday can quietly hurt today, and the only way to know is to keep the measuring stick close and use it every time you touch the query layer.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

A rewrite replaces corpus-matching terms with “yearly charge” when the corpus uses “annual fee.” The relevant document drops out of the top-k. What is the most useful first diagnosis?
Question 1 of 2Debugging

Focus: Diagnose a harmful rewrite by connecting changed tokens to retrieval results and checking against a direct baseline.

A rewrite improves retrieval at k=3 with one retriever. Which follow-up best tests whether that result is robust rather than tied to this setup?
Question 2 of 2Scenario Interpretation

Focus: Design a follow-up experiment that tests whether a rewrite's apparent gain generalizes beyond one retrieval setup.

References

  1. [PDF] RaFe: Ranking Feedback Improves Query Rewriting for RAGaclanthology.org
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.