Skip to content
intermediate

Build and Debug a Small RAG Pipeline

A clean pipeline that answers confidently and wrongly is not a mystery. It is a trace you have not read yet.

Published 2026-10-03Updated 2026-10-048 min read
Hand dipping an egg into blue dye, surrounded by painting supplies, for Easter preparations.
Hand dipping an egg into blue dye, surrounded by painting supplies, for Easter preparations. Photo by Andy Barbour on Pexels.

A clean pipeline that answers confidently and wrongly is not a mystery. It is a trace you have not read yet.

You already know the stages: load, chunk, embed, store, retrieve, assemble, generate. You know retrieval quality should be judged separately from the generated answer. What you have not done yet is wire those stages into a loop you can run, break, and repair.

That is the whole job here. By the end you will have a small working pipeline, a tiny evaluation table, one visible failure, and a measured fix. Not a chatbot. A test bench.

What You Are Actually Building

Scope this down hard. One corpus of a handful of short documents. One embedding model. One vector store. One generation call. Nothing else.

The deliverable is two things:

  • A script that runs end to end and prints what it retrieved before it prints what it answered.
  • A small evaluation set of questions with known correct source chunks.

Environment assumptions, stated plainly: Python 3.10+, an embedding model you can call locally or through an API, an in-memory index, and a generation model with whatever credentials it needs. If you have been following along with the earlier pipeline and retrieval-evaluation material, this is where those two ideas meet: the stages become code, and the evaluation criteria become a table.

Name your success criteria before you write a line:

CriterionQuestion it answers
Retrieval hitIs the correct chunk in the top-k results?
GroundingDoes the answer use only retrieved text?

Two criteria, scored separately. That separation is the entire diagnostic engine.

The Smallest Runnable Pipeline

Get to a visible run fast. Each stage should print something you can inspect. The example below uses a deterministic local embedding so you can run it without any API key and see real numbers. Swap in a real embedding model later; the debugging logic does not change.

import math, re

docs = [
    {"id": "refunds",  "text": "Refunds are issued within 14 days of purchase. "
                               "Digital goods are non-refundable after download."},
    {"id": "shipping", "text": "Standard shipping takes 5-7 business days. "
                               "Express shipping takes 1-2 business days."},
    {"id": "warranty", "text": "Hardware carries a 12-month warranty covering "
                               "manufacturing defects, not accidental damage."},
]

def chunk(doc):
    parts = [s.strip() for s in re.split(r"(?<=[.!?])\s+", doc["text"]) if s.strip()]
    return [{"doc_id": doc["id"], "chunk_id": f'{doc["id"]}-{i}', "text": p}
            for i, p in enumerate(parts)]

chunks = [c for d in docs for c in chunk(d)]

# Deterministic stand-in embedding: bag-of-words hashed into a fixed vector.
# Same function for docs and queries. That is the only rule that matters here.
def embed(text, dim=64):
    vec = [0.0] * dim
    for tok in re.findall(r"[a-z0-9]+", text.lower()):
        vec[hash(tok) % dim] += 1.0
    norm = math.sqrt(sum(v * v for v in vec)) or 1.0
    return [v / norm for v in vec]

def cosine(a, b):
    return sum(x * y for x, y in zip(a, b))

index = [(c, embed(c["text"])) for c in chunks]

def retrieve(query, k=3):
    q = embed(query)
    scored = [(c, cosine(q, v)) for c, v in index]
    return sorted(scored, key=lambda x: x[1], reverse=True)[:k]

def generate(prompt):
    # Stand-in for your LLM call. Return the prompt so you can inspect it.
    return f"[MODEL WOULD ANSWER FROM]\n{prompt}"

def answer(query, k=3):
    hits = retrieve(query, k)
    print("RETRIEVED:", [(c["chunk_id"], round(s, 3)) for c, s in hits])
    context = "\n".join(c["text"] for c, _ in hits)
    return generate(f"Answer using only this context:\n{context}\n\nQ: {query}")

print(answer("How long do refunds take?"))

Run it. You will see two things: a list of retrieved chunk IDs with similarity scores, then the assembled prompt. That is your trace.

Why it behaves this way: the query is embedded with the same function as the documents, so both live in the same vector space and cosine similarity is meaningful. Retrieval ranking decides what the model ever sees. The generation call cannot use evidence that never reached the prompt. That single sentence is the root of every debugging decision that follows.

Note: Print the retrieved chunk IDs and scores, not just the answer. The retrieval trace is your debugging surface. The answer is a symptom.

Knowledge check

Check your understanding

Answer this question before you continue.

Why does the example use the same embedding function for document chunks and queries?
Single Choice

Focus: Explain why a small retrieval pipeline must embed documents and queries with the same function.

Build a Tiny Evaluation Set Before You Debug

You cannot debug what you cannot measure. Write 5–10 questions whose answers live in exactly one chunk, plus one or two questions the corpus genuinely cannot answer.

eval_set = [
    {"q": "How long do refunds take?",     "expected": "refunds-0"},
    {"q": "How fast is express shipping?", "expected": "shipping-1"},
    {"q": "What does the warranty cover?", "expected": "warranty-0"},
    {"q": "Do you ship to Antarctica?",    "expected": None},  # unanswerable
]

def score(k=3):
    rows = []
    for item in eval_set:
        hits = [c["chunk_id"] for c, _ in retrieve(item["q"], k)]
        hit = item["expected"] in hits if item["expected"] else None
        rows.append((item["q"], item["expected"], hits, hit))
    return rows

for row in score():
    print(row)

Record the expected source chunk for each question. That is your ground truth for retrieval, independent of whatever the model says. Then score two columns:

  • Retrieval hit: is expected in the retrieved set?
  • Grounding pass: is the answer supported by the retrieved text?

The unanswerable questions are not filler. A pipeline that invents a shipping policy for Antarctica has a grounding failure, not a retrieval failure. Those are different bugs with different fixes, and this table is the anchor criterion for everything that follows.

Knowledge check

Check your understanding

Answer this question before you continue.

In the evaluation set, “Do you ship to Antarctica?” has expected source `None`. If the pipeline invents a shipping policy, which diagnosis fits the article's criteria?
Scenario Interpretation

Focus: Distinguish retrieval ground truth from grounding evaluation for an unanswerable question.

Read the Trace: Retrieval Failure vs Grounding Failure

A question flows to retrieved chunks, then branches on whether the correct chunk is present. If absent, the issue is retrieval; if present, check the answer's support in the context to identify a grounding failure or a grounded answer.
Check whether the evidence arrived before deciding whether the answer is at fault.

Here is the split that stops you from blaming the model for a retrieval problem.

FailureWhat the trace showsRoot cause
Retrieval failureCorrect chunk never appears in top-kThe model was never given the evidence
Grounding failureCorrect chunk retrieved, answer ignores or contradicts itThe model added unsupported detail
Quiet third caseChunk retrieved but truncated or splitEvidence arrived incomplete

Walk one failing question. Ask "How long do refunds take?" and suppose the trace returns shipping-0, shipping-1, warranty-0 — no refunds-0 at all. The model then answers something plausible about shipping timelines. That is not a hallucination problem. That is a retrieval problem wearing a hallucination costume.

The decision rule: read the retrieved set before you read the answer. If the right chunk is missing, no prompt rewrite will save you. If the right chunk is present and the answer still drifts, now you have a grounding problem worth fixing.

Knowledge check

Check your understanding

Answer this question before you continue.

For “How long do refunds take?”, the trace returns `shipping-0`, `shipping-1`, and `warranty-0`, but not `refunds-0`. The answer discusses shipping timelines. What should you diagnose first?
Debugging

Focus: Diagnose a missing-evidence failure from a retrieval trace before changing generation behavior.

Fix One Thing at a Time

Change one lever, rerun the evaluation set, compare against the baseline table. Two simultaneous changes make the result unreadable.

Retrieval-side levers:

  • Chunk size and overlap — changes what a single chunk can contain.
  • Top-k — changes how much evidence reaches the prompt.
  • Query phrasing — changes what the embedding matches against.

Grounding-side levers:

  • Instruct the model to answer only from the provided context.
  • Instruct it to say "I don't know" when the context is insufficient.

Two mistakes eat most beginners here.

Common mistake: Raising top-k to fix a grounding failure. It adds more chunks, more noise, and more chances to drift. It does not fix the cause.

Common mistake: Rewriting the prompt to fix a retrieval failure. The model cannot cite evidence it never received. Fix retrieval first.

A Bounded Improvement, Measured

Take the single failing question from the trace and pull the matching lever only. For the refunds question, the correct chunk was missing entirely, so this is retrieval-side. Suppose the refund answer is split across two sentences and top-k was too small to catch the second. Raise k from 3 to 5 and rerun.

Before and after, same question:

BEFORE: [('shipping-0', 0.71), ('shipping-1', 0.68), ('warranty-0', 0.64)]
AFTER:  [('refunds-0', 0.79), ('refunds-1', 0.75), ('shipping-0', 0.71), ...]

Now report the change in the table:

MetricBeforeAfter
Retrieval hits2 / 3 answerable3 / 3 answerable
Grounding passes2 / 43 / 4

Interpret honestly. If raising k helped the refunds question but pushed an irrelevant chunk into the Antarctica answer, that is a tradeoff, not a clean win. Say so. And note the limits: a corpus this small does not predict behavior at scale, and an evaluation set this size is a sanity check, not a benchmark. The point is not the number. The point is that you can now see the number move.

Knowledge check

Check your understanding

Answer this question before you continue.

After increasing k, retrieval hits rise from 2/3 to 3/3 answerable questions and grounding passes rise from 2/4 to 3/4. The Antarctica answer now includes an irrelevant chunk. How should this result be reported?
Comparison Reasoning

Focus: Interpret metric changes and tradeoffs after a bounded retrieval adjustment.

One Experiment to Run Next

Keep the pipeline fixed and change only the chunking boundary. Split on paragraphs instead of sentences, rerun the same evaluation set, and predict the outcome before you run it: which questions should improve, which should degrade, and why.

Then compare your prediction to the observed trace. The gap between what you expected and what the chunks actually did is the real lesson — and it turns this script into a reusable test bench for every retrieval and grounding change you make after it.

Inspect the evidence before you judge the answer. That rule survives long after this corpus does.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

A trace shows the correct chunk in the retrieved set, but the answer adds a detail that the chunk does not support. Which bounded next change best matches the diagnosed failure?
Question 1 of 2Scenario Interpretation

Focus: Select a repair lever that matches a grounding failure rather than a retrieval failure.

To test paragraph-based rather than sentence-based chunking, which procedure best lets you learn from the result?
Question 2 of 2Comparison Reasoning

Focus: Describe how to run and interpret the article's proposed chunking experiment.

References

  1. RAGProbe: An Automated Approach for Evaluating RAG Applicationsarxiv.org
  2. Building agentsdevelopers.openai.com
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.