Skip to content
intermediate

Hybrid Retrieval Math: Combining Lexical and Vector Search

Two retrievers return two different top-5 lists for the same query. You merge them, and the merged order looks like it was decided by a coin flip. It…

Published 2026-10-03Updated 2026-10-0410 min read
Expansive sand dunes under a clear blue sky at Patara Beach, Turkey, showcasing natural beauty.
Expansive sand dunes under a clear blue sky at Patara Beach, Turkey, showcasing natural beauty. Photo by Arthur Shuraev on Pexels.

Two retrievers return two different top-5 lists for the same query. You merge them, and the merged order looks like it was decided by a coin flip. It wasn't. It was decided by arithmetic — you just haven't stated the rule yet.

That's the whole problem with hybrid retrieval. Combining lexical and vector search is not a matter of averaging two scores and hoping for the best. The two retrievers produce different kinds of evidence on incompatible scales, and the moment you merge them, you are making a claim about how much each kind of evidence should count. This article makes that claim explicit, works through the arithmetic by hand, and shows you exactly which assumption moved a document up or down.

If you already know that embeddings map meaning to vectors and that RAG retrieves evidence before generation, you have everything you need. If those two ideas are still fuzzy, read up on them first — this article assumes them and moves on.

Two Retrievers, Two Kinds of Evidence

Lexical retrieval and dense retrieval are not two flavors of the same score. They measure different things, and they fail in opposite directions.

Lexical retrieval — think BM25 and its relatives — scores documents on term and document frequency. It rewards exact surface overlap. If your query contains a rare identifier, an error code, an abbreviation, or domain jargon, lexical search finds the document that contains that string and ranks it high. It is precise, interpretable, and fast. It is also blind to vocabulary mismatch: search for "automobile" and it will not find the document that only says "car."

Dense retrieval scores on vector proximity in an embedding space. It rewards conceptual similarity, so it survives paraphrase. Ask "how do I make my database faster?" and it can surface a passage about query optimization that shares almost no words with your question. It is also confidently wrong about strings it has never seen: ask for PROD-SKU-7842X and it may return PROD-SKU-7842Y as the nearest neighbor, with a high similarity score and no awareness that it answered a different question.

The failure modes are structural, not accidental. Dense models were trained to generalize across language, so exact string matching is what they sacrifice. Sparse models were built for exact retrieval, so semantic generalization is what they never had. That opposition is the entire reason to combine them — and the entire reason the combination is not trivial.

Picture two columns of candidates for one query. The lexical column is keyed on a rare term; the dense column is keyed on a paraphrase. Some documents appear in both columns. The overlap region is where the interesting decisions happen.

Knowledge check

Check your understanding

Answer this question before you continue.

A query contains a rare product identifier, while another query asks for a concept using words absent from the relevant passage. Which pairing best describes the signal each retriever is designed to contribute?
Scenario Interpretation

Focus: Distinguish lexical exact-term evidence from dense semantic-similarity evidence in a retrieval scenario.

Why You Cannot Just Add the Scores

The intuitive move is to add the two scores. It fails for a boring, mechanical reason: the scores don't share a scale.

BM25 scores are unbounded and corpus-dependent. A term that appears in three documents out of a million produces a large score; the same term in a small corpus produces a small one. Cosine similarity is typically bounded between -1 and 1 and depends on the embedding model. Add them directly and whichever scale happens to be wider silently dominates the ranking. You didn't choose a weighting. The units chose it for you.

You can fix this with normalization. Min-max scaling or z-scoring brings both scores into a comparable range, and then a weighted sum works. But the normalization is computed over the candidate set, which means the same document gets a different normalized score depending on who else was retrieved. Add a third candidate and the first document's normalized value shifts.

DocumentRaw BM25Raw cosineMin-max normalized BM25 (2 candidates)Min-max normalized BM25 (3 candidates)
A12.00.811.000.80
B4.00.790.000.00
C—0.77—1.00

Document A didn't change. Its score did, because the candidate pool changed.

Rank-based fusion sidesteps the scale problem entirely by throwing away magnitude and keeping only position. That is also its cost: a document that barely won rank 1 and one that won by a mile look identical to the fusion rule.

Note: Fusion assumes each retriever's ordering is informative enough that position carries the signal, and that the two orderings are not so correlated that combining them adds nothing. If both retrievers return nearly the same list, you are paying for two indexes to get one ranking.

Knowledge check

Check your understanding

Answer this question before you continue.

Document A's raw BM25 score stays at 12.0. Its min-max normalized BM25 value is 1.00 with candidates A and B, but 0.80 after candidate C is added. What explains this change?
Misconception Check

Focus: Explain why candidate-set normalization can change a document's normalized score even when its raw score is unchanged.

Reciprocal Rank Fusion, Stated Precisely

Reciprocal rank fusion (RRF) is the most common rank-based method because it needs no labeled data and no calibration. Here is the notation, then the rule.

For a document dd and a retriever rr, let rankr(d)\text{rank}_r(d) be dd's 1-based position in rr's result list. Let kk be a smoothing constant. Then:

RRF(d)=∑r1k+rankr(d)\text{RRF}(d) = \sum_{r} \frac{1}{k + \text{rank}_r(d)}

Documents absent from a retriever's list contribute nothing from that retriever. Read the formula as a mechanism, not a magic number. The kk constant flattens the curve so that the gap between rank 1 and rank 2 matters less than the gap between rank 1 and rank 50. Small kk makes the top of each list dominate; large kk pushes the fusion toward a simple vote where every retrieved document counts almost equally.

Three assumptions are baked in, and you should say them out loud before trusting the output:

  1. Ranks are comparable across retrievers. Rank 3 in the lexical list means roughly the same thing as rank 3 in the dense list.
  2. Each retriever returns a comparable depth of candidates. If one returns 10 and the other returns 100, the deeper list contributes reciprocal scores to far more documents.
  3. No retriever is systematically better in a way equal weighting ignores. RRF weights both retrievers equally by default.

Common mistake: Treating RRF as the best fusion function. It is popular because it is parameter-light and needs no training data — not because it is optimal. Research on fusion functions has found RRF to be sensitive to its parameter, and that a tuned convex combination of normalized scores can outperform it in both in-domain and out-of-domain settings, often with only a small set of labeled examples. RRF is a strong default. It is not a guarantee.

Knowledge check

Check your understanding

Answer this question before you continue.

With RRF smoothing constant k = 60, what contribution does a document at rank 3 in one retriever's list receive from that list?
Single Choice

Focus: Calculate a document's single-list reciprocal-rank contribution using the stated RRF rule.

Worked Example: Fusing Two Ranked Lists

A compact matrix for documents A through E shows lexical ranks 1, 2, 3, 4, 5; dense ranks 2, 4, 1, 5, 3; RRF scores 0.03252, 0.03176, 0.03226, 0.03101, 0.03125; and fused positions 1, 3, 2, 5, 4, respectively.
RRF adds each document’s reciprocal-rank contributions; the resulting order rewards strong positions across both lists.

Five documents, two rankings, k=60k = 60 to match the common default. The lists disagree in the middle, which is where fusion earns its keep.

Lexical ranking: A, B, C, D, E Dense ranking: C, A, E, B, D

Compute each document's reciprocal contribution from each list and sum.

For document A: lexical rank 1 gives 1/(60+1)=0.016391/(60+1) = 0.01639; dense rank 2 gives 1/(60+2)=0.016131/(60+2) = 0.01613. Total: 0.03252.

For document C: lexical rank 3 gives 1/(60+3)=0.015871/(60+3) = 0.01587; dense rank 1 gives 1/(60+1)=0.016391/(60+1) = 0.01639. Total: 0.03226.

For document B: lexical rank 2 gives 0.016130.01613; dense rank 4 gives 1/(60+4)=0.015631/(60+4) = 0.01563. Total: 0.03176.

For document E: lexical rank 5 gives 1/(60+5)=0.015381/(60+5) = 0.01538; dense rank 3 gives 0.015870.01587. Total: 0.03125.

For document D: lexical rank 4 gives 0.015630.01563; dense rank 5 gives 0.015380.01538. Total: 0.03101.

Fused order: A, C, B, E, D.

Look at what happened. A wins because it accumulates the largest total of reciprocal-rank contributions — rank 1 and rank 2 — not because it dominated either list. C won the dense list outright but only placed third lexically, so it lands second. The document that topped a single retriever did not automatically top the fusion. What RRF rewards is not agreement for its own sake; it is the sum of each document's positions across every list it appears in. A document that ranks well in both lists accumulates more than a document that ranks first in one and poorly in the other.

Now re-run the same lists with k=1k = 1. The reciprocal values spread out dramatically: A gets 1/2+1/3=0.8331/2 + 1/3 = 0.833, C gets 1/4+1/2=0.7501/4 + 1/2 = 0.750, B gets 1/3+1/5=0.5331/3 + 1/5 = 0.533. The order holds here, but the margins explode — A's lead over C grows from 0.00026 to 0.083. That steepness is the formula talking: with small kk, the denominator barely grows as rank increases, so the top positions contribute almost all the score. With a different pair of lists, that same sensitivity flips positions.

Knowledge check

Check your understanding

Answer this question before you continue.

Using RRF with k = 60, what is the fused order for the two rankings below?
Output Prediction

Focus: Use the article's two ranked lists and RRF calculation to identify the fused order.

Lexical: A, B, C, D, E
Dense: C, A, E, B, D

What Moves a Document Up or Down

The arithmetic translates directly into retrieval behavior you can predict and debug.

Exact-term queries. A rare identifier or error code gives the lexical retriever a decisive signal, and fusion usually preserves it. The exception: if the dense list is deep enough, many mediocre semantic matches can accumulate rank credit and crowd out the exact hit. This is why candidate depth is not a neutral setting.

Paraphrase queries. The dense retriever carries the ranking. Lexical contributions become noise unless the query happens to share common words with the documents — and then those shared words are often stopwords that inflate lexical scores for the wrong reasons.

Candidate depth is a hidden lever. Retrieving 10 documents from one retriever and 100 from the other quietly reweights the fusion, because the deeper list contributes reciprocal scores to more documents. If you didn't choose that asymmetry deliberately, it chose you.

Correlated retrievers. If both lists are nearly identical, fusion adds cost and latency without adding evidence. Check the overlap before you assume hybrid is buying you anything.

Warning: A higher fused rank means the two retrievers agreed more. It does not mean the document is more likely to answer the question. Fusion reorders candidates. It does not verify relevance.

When Hybrid Fusion Earns Its Cost

Use fusion when queries genuinely mix exact-match intent and conceptual intent, when the corpus contains both jargon and paraphrase, and when you have no labeled data to tune a weighted score combination. Those are the conditions where RRF's parameter-light default is worth its overhead.

Skip it when one retriever already dominates on your query distribution, when latency and index size are the binding constraint, or when a reranking model downstream will reorder the top candidates anyway. Running two retrievers to feed a cross-encoder that ignores half the signal is wasted work.

The practical pattern most teams land on: fuse to widen recall, then rerank a shortlist with a cross-encoder or a small labeled set to recover precision. Fusion is a candidate generator, not the final judge. And measure the fused ranking against each single-retriever baseline on your own queries — retrieval performance does not transfer cleanly across datasets, so a method that wins on someone else's benchmark may lose on yours.

Before you trust any fused ranking, write down three things: which fusion rule you used, what kk or weighting you chose, and how deep each retriever went. Then pull one query from your own retrieval logs where the two retrievers disagree, hand-compute the fused order, and compare it to what your pipeline returned. If the arithmetic doesn't match, you've found a bug. If it does match and the order still looks wrong, you've found a fusion rule that needs a different assumption — and now you know which one.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

A document rises in a fused ranking because both retrievers place it relatively high. What conclusion is justified by the fused rank alone?
Question 1 of 2Misconception Check

Focus: Explain why a high fused rank indicates agreement in positions rather than verified relevance.

A corpus contains domain jargon and paraphrased descriptions, queries mix exact-term and conceptual intent, and there is no labeled data available to tune a weighted score combination. Which strategy best fits the article's guidance?
Question 2 of 2Scenario Interpretation

Focus: Select conditions under which hybrid fusion is a useful candidate-generation strategy.

References

  1. Paper page - An Analysis of Fusion Functions for Hybrid Retrievalhuggingface.co
  2. What is hybrid search? How it works and when to use it | Elasticwww.elastic.co
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.