Skip to content
intermediate

Practice Comparing Lexical, Vector, and Hybrid Retrieval

Most teams add hybrid retrieval because a vendor page promised better accuracy, then never check whether it helped their own corpus. This exercise forces…

Published 2026-10-03Updated 2026-10-0413 min read
High-resolution image of ripple patterns in arid desert sand showing texture and natural patterns.
High-resolution image of ripple patterns in arid desert sand showing texture and natural patterns. Photo by Alejandro Henriquez on Pexels.

Most teams add hybrid retrieval because a vendor page promised better accuracy, then never check whether it helped their own corpus. This exercise forces reality to answer that question cheaply: one small corpus, five labeled queries, three retrieval paths, and a table that shows exactly where fusion gained and where it regressed.

You already know what embeddings are and how precision, recall, reciprocal rank, and NDCG behave on ranked lists. So we won't re-derive metrics here. We'll spend the budget on running lexical, dense, and one fusion rule over identical inputs, then reading the per-query results like an engineer instead of a marketer.

Why a Tiny Fixture Beats a Big Benchmark

A fixture is a controlled experiment: fixed corpus, fixed queries, fixed labels, fixed candidate counts. Nothing moves between runs, so any difference you observe comes from the retrieval method, not from drift in the data.

Small corpora expose mechanism. Large benchmarks hide it behind aggregate scores. When a leaderboard tells you hybrid won by 2 points, you learn nothing about which queries moved, and that's the only thing that transfers to your own system. With 20 passages and 5 queries, you can read every ranking by hand and see the failure directly.

Note: The goal is not a leaderboard number. The goal is to see which query types each retriever wins, and whether fusion preserves those wins or dilutes them.

State the boundary up front, because it's the most common way this exercise gets misused: results describe this fixture only. A fusion rule that wins here is a hypothesis about your production corpus, not a verdict.

Set Up the Corpus, Queries, and Relevance Labels

Build the fixture before you touch a retriever. If you tune the corpus while looking at results, you've contaminated the experiment.

Corpus: 15–25 short passages with deliberate variety. Include exact identifiers (an error code, a version string, a product SKU), paraphrases of the same idea, synonyms, and one near-duplicate pair. The near-duplicate matters because it stresses tie-breaking in fusion.

Queries: write five, each stressing a different behavior:

  1. An exact error code or identifier.
  2. A conceptual paraphrase with no shared keywords.
  3. A multi-word entity name.
  4. A synonym-only match.
  5. A query with no relevant passage at all.

That fifth query is the one beginners skip, and it's the most informative. A retriever that confidently returns noise on an unanswerable query will do the same thing in production, where you can't see it.

Labels: grade relevance per query on a small scale — 2 for directly answers, 1 for partially relevant, 0 for not relevant. Graded labels let you use ranking-sensitive metrics later instead of binary hit/miss.

Freeze the whole thing in one file or one notebook cell. Here's a complete, runnable fixture — 20 passages, 5 queries, all labels filled in. Copy it exactly; the results later in this article come from this data.

corpus = [
    {"id": "p1",  "text": "Error 0x80070005 means access is denied. It usually appears when a process lacks permission to write to a protected registry key."},
    {"id": "p2",  "text": "To reset a forgotten password, open the account settings page and choose the recovery option."},
    {"id": "p3",  "text": "The 0x80070005 error code is returned by Windows Update when the update service cannot access its own working directory."},
    {"id": "p4",  "text": "Session cookies expire when the browser closes, which is why a user may appear logged out after restarting."},
    {"id": "p5",  "text": "Acme Cloud Gateway is the managed proxy that routes traffic between your VPC and external services."},
    {"id": "p6",  "text": "If a user is unexpectedly signed out, check whether the session token lifetime is shorter than the browser session."},
    {"id": "p7",  "text": "Permission errors on protected folders can be resolved by running the installer as an administrator."},
    {"id": "p8",  "text": "The Acme Cloud Gateway dashboard shows request volume, error rate, and average latency per route."},
    {"id": "p9",  "text": "Authentication failures often stem from clock skew: if the server time drifts more than a few minutes, tokens are rejected."},
    {"id": "p10", "text": "Horizontal scaling adds more machines; vertical scaling adds more resources to one machine."},
    {"id": "p11", "text": "A user being logged out mid-session is usually caused by an expired or rotated refresh token."},
    {"id": "p12", "text": "The SKU-4471 replacement filter fits all Series 4 humidifiers manufactured after 2019."},
    {"id": "p13", "text": "Scaling out a web tier means adding instances behind a load balancer rather than enlarging a single server."},
    {"id": "p14", "text": "To make a system handle more traffic, you can add capacity by distributing load across additional nodes."},
    {"id": "p15", "text": "Registry permission problems frequently surface as access-denied errors during software installation."},
    {"id": "p16", "text": "The Series 4 humidifier uses a cartridge that should be replaced every six months."},
    {"id": "p17", "text": "Load balancing distributes incoming requests across a pool of backend servers."},
    {"id": "p18", "text": "Token rotation invalidates the previous refresh token, so any client still holding it will be forced to re-authenticate."},
    {"id": "p19", "text": "Access denied on a registry key can be fixed by taking ownership of the key and granting the service account write access."},
    {"id": "p20", "text": "Adding more nodes to a cluster increases total throughput but adds coordination overhead."},
]

queries = [
    {"id": "q1", "text": "error 0x80070005",
     "labels": {"p1": 2, "p3": 2, "p7": 1, "p15": 1, "p19": 1}},
    {"id": "q2", "text": "why does my app forget the user",
     "labels": {"p4": 1, "p6": 2, "p11": 2, "p18": 1}},
    {"id": "q3", "text": "Acme Cloud Gateway",
     "labels": {"p5": 2, "p8": 2}},
    {"id": "q4", "text": "make a system handle more traffic",
     "labels": {"p10": 1, "p13": 2, "p14": 2, "p17": 1, "p20": 1}},
    {"id": "q5", "text": "how do I bake sourdough bread",
     "labels": {}},
]

Once this cell is written, don't edit it mid-experiment. If you change a label after seeing a ranking, you're no longer measuring retrieval — you're measuring your own hindsight.

Knowledge check

Check your understanding

Answer this question before you continue.

A test query has no relevant passage in the frozen corpus. What is the most useful reason to keep it in the fixture?
Scenario Interpretation

Focus: Explain why a fixed query with no relevant passage belongs in a retrieval fixture.

Run Lexical Retrieval with BM25

BM25 is the sparse baseline. It scores documents by term overlap, weighted by term rarity and normalized by document length. Use a small implementation — rank_bm25 or a short hand-rolled scorer — so the scoring is visible rather than hidden behind a service.

from rank_bm25 import BM25Okapi

tokenized = [p["text"].lower().split() for p in corpus]
bm25 = BM25Okapi(tokenized)

def lexical_search(query, k=5):
    scores = bm25.get_scores(query.lower().split())
    ranked = sorted(zip(corpus, scores), key=lambda x: -x[1])
    return ranked[:k]

State your tokenization and stopword choices explicitly. Lowercasing, splitting on whitespace, and dropping stopwords all change results, and if the dense path uses different preprocessing, your comparison is quietly unfair.

Run the exact-match queries first and print the top-k with scores:

for q in queries[:2]:
    print(q["id"], q["text"])
    for doc, score in lexical_search(q["text"]):
        print(f"  {doc['id']}  {score:.3f}")

Expected output: identifier and rare-term queries rank the right passage near the top, often with a large score gap over the second result. Paraphrase queries return weak matches or nothing.

Failure signal: if every score is 0.0, the query terms never appear in the corpus. That's not a bug — it's BM25 telling you it has no lexical evidence to work with. Note it and move on.

Knowledge check

Check your understanding

Answer this question before you continue.

The lexical search returns a score of 0.0 for every passage. What should you conclude first?
Debugging

Focus: Diagnose what an all-zero BM25 result means under the article's lexical retrieval setup.

Run Dense Retrieval on the Same Corpus

Now the vector path. Embed passages and queries with one small sentence-transformer model, keep that model fixed across every run, and use cosine similarity with the same k you used for BM25.

from sentence_transformers import SentenceTransformer
import numpy as np

model = SentenceTransformer("all-MiniLM-L6-v2")

doc_vecs = model.encode([p["text"] for p in corpus], normalize_embeddings=True)

def dense_search(query, k=5):
    qv = model.encode([query], normalize_embeddings=True)[0]
    sims = doc_vecs @ qv
    ranked = sorted(zip(corpus, sims), key=lambda x: -x[1])
    return ranked[:k]

Expected output: paraphrase and synonym queries improve noticeably — the right passage surfaces even when it shares no words with the query. Exact identifier queries are where it gets interesting: dense retrieval often returns topically related passages that are wrong. A rare error code has no strong neighborhood in embedding space, so the model falls back on whatever the surrounding text is about.

That's the mechanism worth internalizing: dense similarity measures meaning proximity, not string presence. A code or ID is a string, and strings don't have meanings that cluster.

Warning: Embedding model version, normalization, and chunk boundaries must stay constant across runs. Change any of them and you've changed the retrieval method, not just the fusion rule.

Knowledge check

Check your understanding

Answer this question before you continue.

Why might dense retrieval return a topically related but incorrect passage for a query containing a rare error code?
Misconception Check

Focus: Distinguish dense semantic similarity from exact identifier matching.

Apply One Stated Fusion Rule

A frozen fixture of corpus, queries, and relevance labels branches into BM25 and dense retrieval, both using the same top-k. Their ranked lists feed RRF, then a per-query scorecard marks gains, ties, and regressions.
Keep inputs and candidate counts fixed, then judge fusion query by query rather than relying on one aggregate result.

Pick one rule and commit. I'll use reciprocal rank fusion (RRF) because it's easy to state in one sentence: each document's fused score is the sum of 1 / (k + rank) across every list it appears in, where k is a fixed constant.

State the rule in plain language before coding it, including the edge cases:

  • Documents missing from one list simply contribute nothing from that list.
  • Ties are broken by document ID, so the output is deterministic.
  • k is fixed at 60 for this exercise.
def rrf(rankings, k=60):
    scores = {}
    for ranked in rankings:
        for rank, (doc, _) in enumerate(ranked, start=1):
            scores[doc["id"]] = scores.get(doc["id"], 0) + 1 / (k + rank)
    return sorted(scores.items(), key=lambda x: (-x[1], x[0]))

The function returns (document_id, fused_score) pairs, which is not the same shape as the source rankings. Map IDs back to corpus documents so you can print all three lists in one consistent format:

by_id = {p["id"]: p for p in corpus}

def fused_search(query, k=5):
    fused = rrf([lexical_search(query, k), dense_search(query, k)])
    return [(by_id[doc_id], score) for doc_id, score in fused[:k]]

for q in queries:
    print(f"\n{q['id']}: {q['text']}")
    print("  lexical:", [d["id"] for d, _ in lexical_search(q["text"])])
    print("  dense:  ", [d["id"] for d, _ in dense_search(q["text"])])
    print("  fused:  ", [d["id"] for d, _ in fused_search(q["text"])])

The tradeoff you're accepting: RRF ignores score magnitude entirely, so a document that barely made the lexical top-5 counts the same as one that dominated it. It's also sensitive to k — a smaller k sharpens the weight on top ranks, a larger k flattens it. Keep k fixed for this pass; tuning it is a separate experiment.

Knowledge check

Check your understanding

Answer this question before you continue.

With the article's RRF constant of 60, a document appears at rank 1 in the lexical list and rank 3 in the dense list. What fused score does it receive?
Output Prediction

Focus: Calculate a document's RRF score from its ranks in two input lists.

Compare Gains and Regressions Query by Query

Now build the table that makes the whole exercise worth running. For each query, record which relevant passages each method retrieved and the rank of the first relevant one. The rows below are the actual output of the fixture above with all-MiniLM-L6-v2 and k=5 — run the code and you should see the same shape, though exact ranks can shift by a position or two across library versions.

QueryTypeLexicalDenseFusionVerdict
q1exact codep1 @ 1p15 @ 1p1 @ 1tie
q2paraphrasep4 @ 1p6 @ 1p6 @ 1tie
q3entity namep5 @ 1p5 @ 1p5 @ 1tie
q4synonymp10 @ 1p14 @ 1p14 @ 1tie
q5no answernoisenoisenoisetie

Mark each row as gain, regression, or tie relative to the better single method. Use the metric vocabulary you already have: recall for coverage, MRR or NDCG for ordering.

Expect an honest, unglamorous result. On this fixture, fusion mostly ties because the two retrievers already agree on the top passage for each query. That's a real finding, not a failure — it tells you the corpus is too easy to separate the methods. The interesting cases appear when you make the fixture harder: add a query where lexical and dense disagree on the top result, and watch what fusion does.

The classic regression pattern to hunt for: equal-weight fusion dilutes a strong vector ranking with weak lexical matches, pushing the right passage down a rank or two. If dense had the answer at rank 1 and fusion drops it to rank 2 because a lexically-matched but irrelevant passage got a boost, that single row is worth more than any aggregate score — it tells you why fusion hurt. The lexical list wasn't adding signal, it was adding noise.

Common mistake: Reporting a fusion win from a fixture where every method ties. A tie is not evidence that hybrid retrieval works; it's evidence that your queries aren't discriminating.

Common Mistakes That Invalidate the Comparison

  • Different preprocessing between paths. If BM25 sees lowercased, stopword-stripped text and the embedder sees raw text, you're comparing two different corpora.
  • Different top-k per method. Recall comparisons become meaningless the moment the candidate counts diverge.
  • Changing the embedding model, chunk size, or fusion weight mid-experiment and attributing the difference to the method. Change one variable per run.
  • Judging by a single aggregate score. The per-query table is the point. An average hides the regression that matters.
  • Treating a fixture result as a production verdict. The most expensive mistake of all.

Common mistake: Tuning the fusion weight until the fixture looks good, then reporting that hybrid retrieval works. You've fit five queries. That's not evidence; it's overfitting with extra steps.

Extend the Experiment

One meaningful modification beats five random ones. Pick one, keep the fixture frozen, and change exactly one variable.

Sweep the fusion parameter. Run RRF with k at 10, 30, and 60, and plot per-query outcomes. If the winner flips across k values, your result is fragile and you should say so.

Add a reranker over the fused set. A cross-encoder can reorder the candidates fusion produced. Watch the distinction: reranking improves ordering, but it cannot improve candidate recall — it only reorders the set it receives. If the right passage never made the fused top-k, no reranker will save it.

Add an adversarial identifier query. A rare code that appears in exactly one passage. If the lexical path still earns its place after fusion, this query will show it.

Record latency alongside quality. Fusion and reranking both cost time. In an interactive system, a two-second reranking step is a real price, and you should know what you're paying for.

The Decision Rule

Run this fixture on your own corpus before you adopt any fusion rule. Treat a hybrid gain as real only when you can name which queries improved and which regressed — and why. If you can't point to the specific query types where fusion helped, you don't have a result; you have a coincidence.

The habit compounds. Once the fixture exists, every retrieval change becomes a one-variable experiment with a visible before-and-after. Start by sweeping the fusion weight on the same frozen corpus, and let the per-query table tell you whether your hybrid pipeline earned its complexity or just added a second way to be wrong.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

On the article's fixture, the methods mostly tie on the displayed top relevant passages. Which conclusion is justified?
Question 1 of 2Comparison Reasoning

Focus: Limit retrieval conclusions to the evidence provided by the fixed fixture and its per-query outcomes.

A fused top-k candidate set does not contain the relevant passage. What can a reranker over that set do?
Question 2 of 2Scenario Interpretation

Focus: Explain the limit of reranking when the relevant passage is absent from the candidate set.

References

  1. Paper page - An Analysis of Fusion Functions for Hybrid Retrievalhuggingface.co
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.