The Candidate-Recall Ceiling in RAG Reranking
You swap in a stronger reranker. The top three results look sharper, the ordering feels smarter, and the demo lands. Then you measure recall and it barely…

Key topics
You swap in a stronger reranker. The top three results look sharper, the ordering feels smarter, and the demo lands. Then you measure recall and it barely moves. The reranker got better; the answers did not.
That gap is not a tuning problem. It is a structural one, and it has a number.
Why a Better Reranker Stops Helping
The instinct behind the upgrade is reasonable: reranking is a quality stage, so a stronger model should raise quality. That model is incomplete. A reranker does not search. It reorders.
Concretely, a reranker takes a set of candidates that retrieval already produced and returns a permutation of some subset of that set. It is a closed-candidate operator: its output is built entirely from its input. If the relevant document was never retrieved, the reranker never sees it, never scores it, and never ranks it. No cross-encoder, however strong, can promote a document that is not in the pile.
Note: A document the first stage never surfaced cannot be ranked. That single sentence is the whole article in miniature.
This is why reranking is justified only when the candidate set already contains the evidence. The reranker's job is to put the right evidence on top, not to find evidence that retrieval missed. When you treat it as a quality upgrade rather than a reordering of a fixed set, you spend effort on the wrong stage and wonder why the numbers stall.
Notation: Candidate Set, Window, and Two Kinds of Recall
Before the arithmetic, name the pieces. Assume a corpus of documents and a single query with a known set of relevant documents.
- Corpus: every document the system could retrieve.
- Relevant set : the documents that actually answer the query. is how many there are.
- Candidate set : what the first stage returns, of size . This is your recall@K pool.
- Reranking window : the subset of the reranker actually scores. can be smaller than when a context or latency budget caps how many candidates reach the cross-encoder.
Now separate two quantities that get conflated in practice.
Candidate recall measures how much of the relevant set survived retrieval:
This is recall@K. It is a property of your retriever, not your reranker.
Conditional reranker recall measures how much of the retrieved relevant evidence the reranker is even allowed to see:
This is a property of your window size, not your reranker's intelligence.
Picture a funnel: corpus → candidate set → window → final top-k. Relevant documents are marked as they fall through. By the time you reach the reranker, some are already gone, and the funnel shows exactly where.
Warning: The closed-candidate assumption is doing real work here. Systems that re-retrieve mid-generation, fuse multiple sources, or generate candidate documents operate outside this model, and the ceiling below does not bound them.
Deriving the End-to-End Recall Ceiling
End-to-end recall is the fraction of the relevant set that reaches the final output. Build it from the two quantities above.
The relevant documents that survive to the window are . Divide by the full relevant set:
Now factor that same quantity through the two stages:
The left factor is candidate recall. The right factor is conditional reranker recall. The middle term cancels — and that cancellation is the entire point.
Because conditional reranker recall is a fraction, it can never exceed 1. So:
Equality holds only when the window contains every retrieved relevant document — that is, when conditional reranker recall is 1. This is the reranking candidate recall ceiling: the reranker lives under a bound that retrieval set, and it cannot lift that bound no matter how good it gets.
The ceiling binds in two ways. A candidate set that is too small ( too low) drops relevant documents before reranking. A window narrower than drops retrieved relevant documents before the reranker scores them. Both are losses the reranker cannot recover.
What reranking actually does is redistribute recall across ranks. It changes where relevant evidence lands in the final list, not whether it exists in the pipeline. That is a real and valuable job. It is just not the job of raising coverage.
Knowledge check
Check your understanding
Answer this question before you continue.
Worked Example: A Labeled Query with Five Relevant Documents
Numbers make this concrete. Suppose a query has relevant documents in the corpus. The first stage returns candidates, and 4 of the 5 relevant documents are inside that set.
One relevant document is already unreachable. No reranker will ever see it.
Now set the reranking window to . Of the 4 retrieved relevant documents, 3 fall inside the window:
End-to-end recall follows:
Check it against the product: . The relationship holds.
| Stage | Quantity | Value | Evidence lost here |
|---|---|---|---|
| Retrieval | candidate recall | 4/5 = 0.80 | 1 relevant doc never retrieved |
| Window | conditional reranker recall | 3/4 = 0.75 | 1 retrieved relevant doc outside |
| End-to-end | final recall | 3/5 = 0.60 | 2 relevant docs unreachable |
Now watch the ceiling. A perfect reranker over this window caps end-to-end recall at 0.60 — it can promote all 3 relevant documents inside , but it cannot touch the fourth retrieved document outside the window, nor the fifth that retrieval never found. The reranker's quality is irrelevant to that number. The ceiling was decided before it ran.
Knowledge check
Check your understanding
Answer this question before you continue.
Solving for the Candidate Recall You Actually Need
The relationship inverts, which is what makes it useful for planning. Start from a target instead of a guess.
Suppose you need end-to-end recall of 0.90, and you have measured conditional reranker recall at 0.75. Rearrange the product:
Plug in the numbers:
A required candidate recall above 1.0 is impossible. That is not a rounding error — it is proof that the target is unreachable at this window size. No amount of retrieval improvement gets you there, because the window is throwing away a quarter of the retrieved relevant evidence before the reranker scores anything.
The fix path has two levers. Widen the window to raise conditional reranker recall, or improve retrieval to raise candidate recall. Try widening until conditional reranker recall reaches 0.95:
Now the target is reachable, but it demands candidate recall near 0.95 — a real retrieval bar, not a reranker setting.
Common mistake: Tuning the reranker when the required candidate recall already exceeds what your retriever achieves. If retrieval cannot hit the number, no reranker closes the gap.
One caveat on measurement: both quantities need a labeled evaluation set, and small samples make these ratios noisy. A handful of queries will produce candidate recall figures that swing by several points. Treat the arithmetic as a planning tool, not a precise instrument, until your fixture is large enough to trust.
Knowledge check
Check your understanding
Answer this question before you continue.
Diagnosing Which Stage Is the Bottleneck
The arithmetic becomes a diagnostic once you run it in order.
- Measure recall@K at your current .
- Measure conditional reranker recall at your chosen window.
- Compare the product against your target.
The pattern of the two numbers tells you where to work.
| Candidate recall | Conditional reranker recall | Diagnosis | Fix |
|---|---|---|---|
| Low | High | Retrieval is the bottleneck | Embeddings, hybrid search, chunking, query rewriting |
| High | Low | Window is the bottleneck | Widen so more retrieved relevant docs reach the reranker |
| Low | Low | Evidence lost twice | Fix retrieval first — it sets the ceiling |
When both are low, fix retrieval first. It sets the ceiling, so improving it raises the maximum that every downstream stage can reach.
Two mistakes show up again and again. The first is reading a strong NDCG or a clean top-3 as proof that coverage is fine. Ordering metrics can look stable while long-tail recall quietly degrades across query variants — the top of the list stays sharp while the tail rots. The second is widening without checking whether the recall curve has already flattened, which adds reranker cost for near-zero coverage gain.
The decision boundary is the shape of the recall@K curve. When recall is still climbing steeply at your current , widen before you touch the reranker — you are leaving retrievable evidence on the table. When the curve has plateaued, more buys almost nothing, and the move is to change the retrieval method: hybrid search, domain-adapted embeddings, or better chunking.
Knowledge check
Check your understanding
Answer this question before you continue.
What Breaks the Ceiling Argument
The bound is precise, so it is worth stating where it stops applying.
It assumes a closed-candidate reranker. Systems that re-retrieve during generation, fuse multiple retrieval sources, or generate candidate documents are not bounded by this relationship, because their candidate set is not fixed before ranking. The bound is also about coverage, not answer quality: a pipeline can hit its recall ceiling and still produce a poor answer for reasons entirely downstream of retrieval. And recall ceilings measured on one corpus do not transfer automatically — the numbers are properties of your retriever and your labeled set, not universal constants.
The practical next step is to build a small labeled query fixture and plot recall@K across a range of . Find the knee — the point where the curve flattens. That knee tells you where to set your candidate budget and your reranking window, and it tells you whether your next hour belongs with the reranker or upstream with embeddings, hybrid search, and chunking.
Measure candidate recall before you tune the reranker. The reranker lives under a ceiling it did not build, and the levers that raise it are the unglamorous ones upstream.
Knowledge check
Final check
Finish the article by checking the ideas you just learned.
References
Want a more structured LLMOps path?
Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.
Large Language Models Starter Pack
A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.
- 227-page Illustrated PDF edition
- 12 guided LLM engineering chapters
- Visual concept diagrams
- Self-assessment quizzes
- Bonus deep-dive sections
- Prompt design, structured output, context windows & RAG pipelines
- Agents, tool calling, prompt injection, evaluation & application lifecycles
Coming soon


