Skip to content
intermediate

Evaluate a RAG Reranking Stage Against Relevance Judgments

The reranker "feels" better. The top result looks more relevant than it did before. And yet you cannot say whether the stage earned its latency, because…

Published 2026-10-03Updated 2026-10-0411 min read
A lone hiker walking on a vast dune at Sossusvlei, Namibia, against a clear sky.
A lone hiker walking on a vast dune at Sossusvlei, Namibia, against a clear sky. Photo by MINEIA MARTINS on Pexels.

The reranker "feels" better. The top result looks more relevant than it did before. And yet you cannot say whether the stage earned its latency, because "looks better" is not a measurement. Here is the smallest experiment that turns that feeling into a verdict: one query, a handful of candidate passages, labels you write down before you run anything, and a cross-encoder that produces an ordering you can compare against them.

If you already know where reranking sits in the pipeline and why a two-stage design exists, you have what you need. The question now is narrower and more useful: did the relevant passages move up, and did the irrelevant ones move down? That is the anchor criterion for everything below.

Why "It Looks Better" Is Not Evidence

The default beginner move is to run a reranker, glance at the new top result, and declare victory. That judgment is unfalsifiable. It does not survive a second query, and it cannot tell you whether the reranker helped or whether the original retrieval was already fine.

The stronger model is boring and repeatable: fix a small fixture, write relevance judgments first, then let the reranker produce an ordering you compare against those labels. The labels are the ground truth. The reranker is the thing on trial.

Be honest about scope before you start. This exercise can show whether the stage helped this fixture. It cannot establish general retrieval quality, and it cannot tell you what will happen on your production corpus. What it can do is give you a reusable procedure and a decision rule you can re-run every time the retriever, the model, or the corpus changes.

Note: A reranking stage is justified only when its ordering beats the initial ranking against labels you wrote down before you ran it. Everything else is vibes.

Set Up the Fixture and the Labels

The fixture is one query, a small candidate list with an initial rank order, and a relevance label per passage. Keep it to roughly 8–12 passages — enough to have a real ordering, small enough to label by hand in a few minutes.

Write the labels before you run the reranker. If you write them afterward, you are not measuring the reranker; you are confirming what you already saw. That is the single most common way this experiment quietly becomes worthless.

I use a simple graded scheme here: 2 for a passage that directly answers the query, 1 for partial or supporting relevance, 0 for irrelevant. Binary labels are defensible too, but graded labels let you see how far a passage moved, not just whether it crossed a line.

Here is the fixture. The query is: "How does a cross-encoder differ from a bi-encoder in retrieval?"

Initial rankPassage (abbreviated)Label
1"Bi-encoders embed query and document separately, enabling fast approximate search."1
2"Vector databases store embeddings and support nearest-neighbor lookup."0
3"A cross-encoder reads the query and passage together and outputs a relevance score."2
4"Reranking adds latency because each candidate needs its own forward pass."1
5"Chunking strategy affects how much context each passage carries."0
6"Cross-encoders are more accurate but slower than bi-encoders for the same candidate set."2
7"Embeddings map text to vectors so similarity can be computed numerically."0
8"The first-stage retriever decides which candidates the reranker can even see."1

Two passages carry label 2 — ranks 3 and 6. The initial ranking buries one of them at position 6. That is the gap the reranker has to close.

Environment assumptions: Python 3.9+, the sentence-transformers package, and a local cross-encoder model. No API key is required for this path.

pip install sentence-transformers

Knowledge check

Check your understanding

Answer this question before you continue.

A learner labels passages only after inspecting the reranked order. What is the main problem with this procedure?
Scenario Interpretation

Focus: Explain why relevance judgments must be fixed before running the reranker.

Run the Cross-Encoder Reranker

A cross-encoder scores a (query, passage) pair jointly rather than embedding the two separately. That joint read is why it costs one forward pass per candidate — and why candidate count is the latency lever you will feel later.

from sentence_transformers import CrossEncoder

query = "How does a cross-encoder differ from a bi-encoder in retrieval?"

passages = [
    ("Bi-encoders embed query and document separately, enabling fast approximate search.", 1),
    ("Vector databases store embeddings and support nearest-neighbor lookup.", 0),
    ("A cross-encoder reads the query and passage together and outputs a relevance score.", 2),
    ("Reranking adds latency because each candidate needs its own forward pass.", 1),
    ("Chunking strategy affects how much context each passage carries.", 0),
    ("Cross-encoders are more accurate but slower than bi-encoders for the same candidate set.", 2),
    ("Embeddings map text to vectors so similarity can be computed numerically.", 0),
    ("The first-stage retriever decides which candidates the reranker can even see.", 1),
]

model = CrossEncoder("cross-encoder/ms-marco-MiniLM-L-6-v2")

pairs = [(query, text) for text, _ in passages]
scores = model.predict(pairs)

# Pair each passage with its score, then sort descending.
ranked = sorted(
    zip(scores, passages),
    key=lambda x: x[0],
    reverse=True,
)

print(f"{'Init':>4} {'New':>4} {'Score':>7} {'Label':>5}  Passage")
for new_pos, (score, (text, label)) in enumerate(ranked, start=1):
    init_pos = passages.index((text, label)) + 1
    print(f"{init_pos:>4} {new_pos:>4} {score:>7.3f} {label:>5}  {text[:60]}")

The output below is an illustrative example, not a guaranteed result. Your scores and ordering will depend on the model version, the library version, and the exact passage text. What matters is the structure: a table with initial rank, new rank, score, label, and passage.

Init  New   Score Label  Passage
   3    1   6.412     2  A cross-encoder reads the query and passage together...
   6    2   5.887     2  Cross-encoders are more accurate but slower...
   1    3   2.104     1  Bi-encoders embed query and document separately...
   4    4   1.733     1  Reranking adds latency because each candidate...
   8    5   0.921     1  The first-stage retriever decides which candidates...
   2    6  -2.310     0  Vector databases store embeddings...
   5    7  -3.045     0  Chunking strategy affects how much context...
   7    8  -3.612     0  Embeddings map text to vectors...

If your run produces a different ordering, that is not a failure. It is your result. The next section shows how to read whatever table you actually got.

Knowledge check

Check your understanding

Answer this question before you continue.

Given the article's `passages` list of `(text, label)` tuples, what does `model.predict(pairs)` score?
Output Prediction

Focus: Identify the query-passage pairs scored by the cross-encoder in the provided implementation.

pairs = [(query, text) for text, _ in passages]
scores = model.predict(pairs)

Read the Before-and-After Table

A compact before-and-after matrix shows precision at 3 rising from 0.67 to 1.00 after reranking, while mean reciprocal rank remains 1.00 in both rankings.
The reranker improves the top-three mix, but MRR misses that gain because the initial top result was already relevant.

Walk your table one row at a time. The signal you want is a relevant passage climbing. In the illustrative output, the passage at initial rank 6 (label 2) moved to position 2 — that is the reranker doing its job. The passage at initial rank 3 (also label 2) moved to position 1. Both relevant passages are now at the top, where the generation stage will actually read them.

Now the counter-signal. Passage at initial rank 1 (label 1) fell to position 3. It is still above the irrelevant passages, so the fall is harmless here — but on a different fixture, a relevant passage falling below an irrelevant one is exactly the failure you are hunting for.

Eyeballing movement is not enough. Compute two simple measures for both orderings. Use your table, not the illustrative one.

Precision at k asks: of the top k passages, how many are relevant? Using label ≥ 1 as relevant and k = 3:

  • Initial top 3: labels 1, 0, 2 → 2 of 3 relevant → 0.67
  • Reranked top 3: labels 2, 2, 1 → 3 of 3 relevant → 1.00

Mean reciprocal rank (MRR) asks: how high is the first relevant passage? With label ≥ 1 as relevant:

  • Initial: first relevant is at rank 1 → 1.00
  • Reranked: first relevant is at rank 1 → 1.00

Notice what just happened. Precision at 3 improved; MRR did not move, because the initial ranking already had a relevant passage on top. The metric you choose decides which improvement you notice. If you only tracked MRR, you would conclude the reranker did nothing. If you only tracked precision at 3, you would miss that the top slot was already fine.

If your run produced a different ordering, the same logic applies. Compute both metrics on your initial and reranked lists. Then ask: which metric moved, which did not, and does the movement matter for the cutoff your generation stage actually uses?

This is the mechanism showing through: the cross-encoder optimizes pairwise relevance between query and passage. It does not optimize your downstream answer. A better ordering is a proxy for a better answer, not a guarantee of one — the generation stage and the context window still shape the final output.

Common mistake: Concluding "the reranker helped" from a single improved metric. State which metric moved, which did not, and why.

Knowledge check

Check your understanding

Answer this question before you continue.

In the article's illustrative table, precision at 3 rises from 0.67 to 1.00, while MRR remains 1.00. What best explains this result?
Comparison Reasoning

Focus: Interpret why precision at k and MRR can report different changes for the same reranking result.

Failure Modes and Debugging Signals

Scores clustered near each other. If every score lands within a narrow band, the model is probably mismatched to your domain, or your passages are too short to discriminate. Try a different reranker model before concluding reranking is useless.

Nothing reordered. If the reranked order matches the initial order, your first-stage retrieval was already strong on this fixture. That is a finding, not a bug. It means the reranker is not earning its latency here.

Label leakage. If you wrote labels after seeing the reranked order, the comparison is worthless. Redo it. This is the failure that looks like success.

Candidate starvation. A reranker can only reorder what retrieval handed it. A relevant passage that never entered the candidate list cannot be rescued by any reranker. If your fixture's relevant passages are all present, you are testing reranking; if some are missing, you are testing retrieval and blaming the wrong stage.

Latency and cost. The cross-encoder runs one forward pass per candidate. Doubling the candidate list roughly doubles the reranking cost. A reranker that wins on quality can still lose on the operational tradeoff — that is a decision, not a bug.

Knowledge check

Check your understanding

Answer this question before you continue.

A relevant passage is missing from the candidates supplied to the reranker. Which diagnosis matches the article?
Debugging

Focus: Diagnose why a reranker cannot recover a relevant passage absent from its candidate list.

One Modification Worth Running

Change one variable at a time. The highest-value change is adding a second query with its own labels, because a single-query fixture cannot distinguish a real improvement from a lucky reorder.

Before you run it, write down your prediction. Will the reranker help the second query the same way? More? Less? Then run it and compare the prediction to the observed ordering. The gap between what you expected and what the table shows is where the learning lives.

Other single-variable changes worth trying: swap the reranker model and re-run the same fixture; expand the candidate list from 8 to 16 passages and watch whether precision at 3 holds; or change the label scheme from graded to binary and see whether your conclusion survives.

Record each run as a short note: fixture, model, metric, before, after, verdict, and what remains unknown. That note is the reusable asset — the experiment is disposable, the record is not.

What This Fixture Cannot Tell You

Eight passages and one query cannot estimate performance across a real corpus, a real domain, or a real query distribution. The numbers you computed are true of this fixture and nothing else.

The labels were written by one person — you. They encode your relevance standard. A different annotator would disagree on the partial-relevance passages, and that disagreement is itself a measurement problem you have not solved.

Improved ordering is a proxy for better answers, not proof of them. The generation stage, the context window, and the prompt all still shape the final output. You have measured one link in the chain.

Published results showing reranking gains come from specific datasets and specific models. They are context for your experiment, not a substitute for it. Your corpus is not their corpus.

The honest conclusion format is this: this stage helped on this fixture, under these conditions, by these metrics — and here is what would have to be true for it to help in production.

The Decision Rule

This fixture can tell you whether the reranker improved ranking quality under your chosen metrics. It cannot tell you whether to keep the stage in production. That decision requires two more inputs: whether the metric improvement matters at the cutoff your generation stage actually uses, and whether the quality gain justifies the added latency and cost.

So the bounded rule is: keep the reranking stage only when it moves labeled-relevant passages upward on a fixture you wrote before running it, and the metric that improved is the one your downstream task cares about, and the latency cost is acceptable for your use case. Any one of those conditions failing is a reason to reconsider.

Re-run the fixture whenever you change the retriever, the model, or the corpus. The fixture is cheap; the wrong architectural decision is not.

Your next step: extend the fixture to several queries and track precision at k before and after reranking across all of them. One query gives you an anecdote. Five queries with labels give you a pattern. That pattern is what turns "it feels better" into a decision you can defend.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

The reranker improves the measured ranking on the article's one-query fixture. Which conclusion is supported?
Question 1 of 2Misconception Check

Focus: State what a small single-query fixture can and cannot establish.

A reranker moves labeled-relevant passages upward and improves a metric, but that metric is not important at the generation stage's cutoff. Latency is acceptable. Under the article's decision rule, what should you do?
Question 2 of 2Scenario Interpretation

Focus: Apply the article's combined quality, downstream-relevance, and operational-cost decision rule.

References

  1. RAG reranking explained: better context, better answers | Meilisearchwww.meilisearch.com
  2. Rerankers and Two-Stage Retrieval | Pineconewww.pinecone.io
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.