Skip to content
intermediate

The Candidate-Recall Ceiling in RAG Reranking

You swap in a stronger reranker. The top three results look sharper, the ordering feels smarter, and the demo lands. Then you measure recall and it barely…

Published 2026-10-03Updated 2026-10-049 min read
Detailed close-up of sand dunes showcasing texture and patterns in natural light.
Detailed close-up of sand dunes showcasing texture and patterns in natural light. Photo by Ulrick Trappschuh on Pexels.

You swap in a stronger reranker. The top three results look sharper, the ordering feels smarter, and the demo lands. Then you measure recall and it barely moves. The reranker got better; the answers did not.

That gap is not a tuning problem. It is a structural one, and it has a number.

Why a Better Reranker Stops Helping

The instinct behind the upgrade is reasonable: reranking is a quality stage, so a stronger model should raise quality. That model is incomplete. A reranker does not search. It reorders.

Concretely, a reranker takes a set of candidates that retrieval already produced and returns a permutation of some subset of that set. It is a closed-candidate operator: its output is built entirely from its input. If the relevant document was never retrieved, the reranker never sees it, never scores it, and never ranks it. No cross-encoder, however strong, can promote a document that is not in the pile.

Note: A document the first stage never surfaced cannot be ranked. That single sentence is the whole article in miniature.

This is why reranking is justified only when the candidate set already contains the evidence. The reranker's job is to put the right evidence on top, not to find evidence that retrieval missed. When you treat it as a quality upgrade rather than a reordering of a fixed set, you spend effort on the wrong stage and wonder why the numbers stall.

Notation: Candidate Set, Window, and Two Kinds of Recall

Before the arithmetic, name the pieces. Assume a corpus of documents and a single query with a known set of relevant documents.

  • Corpus: every document the system could retrieve.
  • Relevant set RR: the documents that actually answer the query. ∣R∣|R| is how many there are.
  • Candidate set CC: what the first stage returns, of size KK. This is your recall@K pool.
  • Reranking window WW: the subset of CC the reranker actually scores. WW can be smaller than CC when a context or latency budget caps how many candidates reach the cross-encoder.

Now separate two quantities that get conflated in practice.

Candidate recall measures how much of the relevant set survived retrieval:

candidate recall=∣C∩R∣∣R∣\text{candidate recall} = \frac{|C \cap R|}{|R|}

This is recall@K. It is a property of your retriever, not your reranker.

Conditional reranker recall measures how much of the retrieved relevant evidence the reranker is even allowed to see:

conditional reranker recall=∣W∩R∣∣C∩R∣\text{conditional reranker recall} = \frac{|W \cap R|}{|C \cap R|}

This is a property of your window size, not your reranker's intelligence.

Picture a funnel: corpus → candidate set CC → window WW → final top-k. Relevant documents are marked as they fall through. By the time you reach the reranker, some are already gone, and the funnel shows exactly where.

Warning: The closed-candidate assumption is doing real work here. Systems that re-retrieve mid-generation, fuse multiple sources, or generate candidate documents operate outside this model, and the ceiling below does not bound them.

Deriving the End-to-End Recall Ceiling

Five relevant-document markers narrow to four in the candidate set, then three in the reranking window. The stages are labeled candidate recall 4/5, conditional reranker recall 3/4, and end-to-end recall 3/5.
A reranker can reorder only the three relevant documents that reach its window; the earlier losses set the coverage ceiling.

End-to-end recall is the fraction of the relevant set that reaches the final output. Build it from the two quantities above.

The relevant documents that survive to the window are ∣W∩R∣|W \cap R|. Divide by the full relevant set:

end-to-end recall=∣W∩R∣∣R∣\text{end-to-end recall} = \frac{|W \cap R|}{|R|}

Now factor that same quantity through the two stages:

∣W∩R∣∣R∣=∣C∩R∣∣R∣×∣W∩R∣∣C∩R∣\frac{|W \cap R|}{|R|} = \frac{|C \cap R|}{|R|} \times \frac{|W \cap R|}{|C \cap R|}

The left factor is candidate recall. The right factor is conditional reranker recall. The middle term ∣C∩R∣|C \cap R| cancels — and that cancellation is the entire point.

end-to-end recall=candidate recall×conditional reranker recall\text{end-to-end recall} = \text{candidate recall} \times \text{conditional reranker recall}

Because conditional reranker recall is a fraction, it can never exceed 1. So:

end-to-end recall≤candidate recall\text{end-to-end recall} \le \text{candidate recall}

Equality holds only when the window contains every retrieved relevant document — that is, when conditional reranker recall is 1. This is the reranking candidate recall ceiling: the reranker lives under a bound that retrieval set, and it cannot lift that bound no matter how good it gets.

The ceiling binds in two ways. A candidate set that is too small (KK too low) drops relevant documents before reranking. A window WW narrower than CC drops retrieved relevant documents before the reranker scores them. Both are losses the reranker cannot recover.

What reranking actually does is redistribute recall across ranks. It changes where relevant evidence lands in the final list, not whether it exists in the pipeline. That is a real and valuable job. It is just not the job of raising coverage.

Knowledge check

Check your understanding

Answer this question before you continue.

Under the article’s closed-candidate model, when does end-to-end recall equal candidate recall?
Misconception Check

Focus: Identify the condition under which end-to-end recall reaches the candidate-recall ceiling.

Worked Example: A Labeled Query with Five Relevant Documents

Numbers make this concrete. Suppose a query has ∣R∣=5|R| = 5 relevant documents in the corpus. The first stage returns K=20K = 20 candidates, and 4 of the 5 relevant documents are inside that set.

candidate recall=45=0.80\text{candidate recall} = \frac{4}{5} = 0.80

One relevant document is already unreachable. No reranker will ever see it.

Now set the reranking window to W=10W = 10. Of the 4 retrieved relevant documents, 3 fall inside the window:

conditional reranker recall=34=0.75\text{conditional reranker recall} = \frac{3}{4} = 0.75

End-to-end recall follows:

end-to-end recall=35=0.60\text{end-to-end recall} = \frac{3}{5} = 0.60

Check it against the product: 0.80×0.75=0.600.80 \times 0.75 = 0.60. The relationship holds.

StageQuantityValueEvidence lost here
Retrievalcandidate recall4/5 = 0.801 relevant doc never retrieved
Windowconditional reranker recall3/4 = 0.751 retrieved relevant doc outside WW
End-to-endfinal recall3/5 = 0.602 relevant docs unreachable

Now watch the ceiling. A perfect reranker over this window caps end-to-end recall at 0.60 — it can promote all 3 relevant documents inside WW, but it cannot touch the fourth retrieved document outside the window, nor the fifth that retrieval never found. The reranker's quality is irrelevant to that number. The ceiling was decided before it ran.

Knowledge check

Check your understanding

Answer this question before you continue.

A query has 5 relevant documents; retrieval returns 4 of them, and 3 of those 4 are inside the reranking window. What is end-to-end recall?
Output Prediction

Focus: Calculate end-to-end recall from the counts of relevant documents retrieved and admitted to the reranking window.

Solving for the Candidate Recall You Actually Need

The relationship inverts, which is what makes it useful for planning. Start from a target instead of a guess.

Suppose you need end-to-end recall of 0.90, and you have measured conditional reranker recall at 0.75. Rearrange the product:

required candidate recall=target end-to-end recallconditional reranker recall\text{required candidate recall} = \frac{\text{target end-to-end recall}}{\text{conditional reranker recall}}

Plug in the numbers:

required candidate recall=0.900.75=1.20\text{required candidate recall} = \frac{0.90}{0.75} = 1.20

A required candidate recall above 1.0 is impossible. That is not a rounding error — it is proof that the target is unreachable at this window size. No amount of retrieval improvement gets you there, because the window is throwing away a quarter of the retrieved relevant evidence before the reranker scores anything.

The fix path has two levers. Widen the window to raise conditional reranker recall, or improve retrieval to raise candidate recall. Try widening WW until conditional reranker recall reaches 0.95:

required candidate recall=0.900.95≈0.947\text{required candidate recall} = \frac{0.90}{0.95} \approx 0.947

Now the target is reachable, but it demands candidate recall near 0.95 — a real retrieval bar, not a reranker setting.

Common mistake: Tuning the reranker when the required candidate recall already exceeds what your retriever achieves. If retrieval cannot hit the number, no reranker closes the gap.

One caveat on measurement: both quantities need a labeled evaluation set, and small samples make these ratios noisy. A handful of queries will produce candidate recall figures that swing by several points. Treat the arithmetic as a planning tool, not a precise instrument, until your fixture is large enough to trust.

Knowledge check

Check your understanding

Answer this question before you continue.

A system targets end-to-end recall of 0.90, while conditional reranker recall is 0.75. What candidate recall would be required, and what does that imply?
Single Choice

Focus: Solve for required candidate recall and recognize when a target is impossible at a given conditional reranker recall.

Diagnosing Which Stage Is the Bottleneck

The arithmetic becomes a diagnostic once you run it in order.

  1. Measure recall@K at your current KK.
  2. Measure conditional reranker recall at your chosen window.
  3. Compare the product against your target.

The pattern of the two numbers tells you where to work.

Candidate recallConditional reranker recallDiagnosisFix
LowHighRetrieval is the bottleneckEmbeddings, hybrid search, chunking, query rewriting
HighLowWindow is the bottleneckWiden WW so more retrieved relevant docs reach the reranker
LowLowEvidence lost twiceFix retrieval first — it sets the ceiling

When both are low, fix retrieval first. It sets the ceiling, so improving it raises the maximum that every downstream stage can reach.

Two mistakes show up again and again. The first is reading a strong NDCG or a clean top-3 as proof that coverage is fine. Ordering metrics can look stable while long-tail recall quietly degrades across query variants — the top of the list stays sharp while the tail rots. The second is widening KK without checking whether the recall curve has already flattened, which adds reranker cost for near-zero coverage gain.

The decision boundary is the shape of the recall@K curve. When recall is still climbing steeply at your current KK, widen before you touch the reranker — you are leaving retrievable evidence on the table. When the curve has plateaued, more KK buys almost nothing, and the move is to change the retrieval method: hybrid search, domain-adapted embeddings, or better chunking.

Knowledge check

Check your understanding

Answer this question before you continue.

A labeled evaluation shows high candidate recall but low conditional reranker recall. Which stage should be investigated first?
Scenario Interpretation

Focus: Use candidate recall and conditional reranker recall together to locate the pipeline bottleneck.

What Breaks the Ceiling Argument

The bound is precise, so it is worth stating where it stops applying.

It assumes a closed-candidate reranker. Systems that re-retrieve during generation, fuse multiple retrieval sources, or generate candidate documents are not bounded by this relationship, because their candidate set is not fixed before ranking. The bound is also about coverage, not answer quality: a pipeline can hit its recall ceiling and still produce a poor answer for reasons entirely downstream of retrieval. And recall ceilings measured on one corpus do not transfer automatically — the numbers are properties of your retriever and your labeled set, not universal constants.

The practical next step is to build a small labeled query fixture and plot recall@K across a range of KK. Find the knee — the point where the curve flattens. That knee tells you where to set your candidate budget and your reranking window, and it tells you whether your next hour belongs with the reranker or upstream with embeddings, hybrid search, and chunking.

Measure candidate recall before you tune the reranker. The reranker lives under a ceiling it did not build, and the levers that raise it are the unglamorous ones upstream.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

Two systems use the same fixed candidate set and window, but one has a stronger reranker. Which claim about coverage is justified by the article?
Question 1 of 2Comparison Reasoning

Focus: Explain why stronger reranking cannot recover relevant documents absent from its fixed candidate input.

Which system is explicitly outside the article’s fixed-candidate ceiling argument?
Question 2 of 2Misconception Check

Focus: Distinguish a fixed-candidate reranking pipeline from systems that can add candidates after initial retrieval.

References

  1. The Recall Ceiling of LLM Recommendation Rerankingarxiv.org
  2. First-pass retrieval: the recall-optimized retrieval stagezeroentropy.dev
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.