Build and Debug a Small RAG Pipeline
A clean pipeline that answers confidently and wrongly is not a mystery. It is a trace you have not read yet.

Key topics
A clean pipeline that answers confidently and wrongly is not a mystery. It is a trace you have not read yet.
You already know the stages: load, chunk, embed, store, retrieve, assemble, generate. You know retrieval quality should be judged separately from the generated answer. What you have not done yet is wire those stages into a loop you can run, break, and repair.
That is the whole job here. By the end you will have a small working pipeline, a tiny evaluation table, one visible failure, and a measured fix. Not a chatbot. A test bench.
What You Are Actually Building
Scope this down hard. One corpus of a handful of short documents. One embedding model. One vector store. One generation call. Nothing else.
The deliverable is two things:
- A script that runs end to end and prints what it retrieved before it prints what it answered.
- A small evaluation set of questions with known correct source chunks.
Environment assumptions, stated plainly: Python 3.10+, an embedding model you can call locally or through an API, an in-memory index, and a generation model with whatever credentials it needs. If you have been following along with the earlier pipeline and retrieval-evaluation material, this is where those two ideas meet: the stages become code, and the evaluation criteria become a table.
Name your success criteria before you write a line:
| Criterion | Question it answers |
|---|---|
| Retrieval hit | Is the correct chunk in the top-k results? |
| Grounding | Does the answer use only retrieved text? |
Two criteria, scored separately. That separation is the entire diagnostic engine.
The Smallest Runnable Pipeline
Get to a visible run fast. Each stage should print something you can inspect. The example below uses a deterministic local embedding so you can run it without any API key and see real numbers. Swap in a real embedding model later; the debugging logic does not change.
import math, re
docs = [
{"id": "refunds", "text": "Refunds are issued within 14 days of purchase. "
"Digital goods are non-refundable after download."},
{"id": "shipping", "text": "Standard shipping takes 5-7 business days. "
"Express shipping takes 1-2 business days."},
{"id": "warranty", "text": "Hardware carries a 12-month warranty covering "
"manufacturing defects, not accidental damage."},
]
def chunk(doc):
parts = [s.strip() for s in re.split(r"(?<=[.!?])\s+", doc["text"]) if s.strip()]
return [{"doc_id": doc["id"], "chunk_id": f'{doc["id"]}-{i}', "text": p}
for i, p in enumerate(parts)]
chunks = [c for d in docs for c in chunk(d)]
# Deterministic stand-in embedding: bag-of-words hashed into a fixed vector.
# Same function for docs and queries. That is the only rule that matters here.
def embed(text, dim=64):
vec = [0.0] * dim
for tok in re.findall(r"[a-z0-9]+", text.lower()):
vec[hash(tok) % dim] += 1.0
norm = math.sqrt(sum(v * v for v in vec)) or 1.0
return [v / norm for v in vec]
def cosine(a, b):
return sum(x * y for x, y in zip(a, b))
index = [(c, embed(c["text"])) for c in chunks]
def retrieve(query, k=3):
q = embed(query)
scored = [(c, cosine(q, v)) for c, v in index]
return sorted(scored, key=lambda x: x[1], reverse=True)[:k]
def generate(prompt):
# Stand-in for your LLM call. Return the prompt so you can inspect it.
return f"[MODEL WOULD ANSWER FROM]\n{prompt}"
def answer(query, k=3):
hits = retrieve(query, k)
print("RETRIEVED:", [(c["chunk_id"], round(s, 3)) for c, s in hits])
context = "\n".join(c["text"] for c, _ in hits)
return generate(f"Answer using only this context:\n{context}\n\nQ: {query}")
print(answer("How long do refunds take?"))
Run it. You will see two things: a list of retrieved chunk IDs with similarity scores, then the assembled prompt. That is your trace.
Why it behaves this way: the query is embedded with the same function as the documents, so both live in the same vector space and cosine similarity is meaningful. Retrieval ranking decides what the model ever sees. The generation call cannot use evidence that never reached the prompt. That single sentence is the root of every debugging decision that follows.
Note: Print the retrieved chunk IDs and scores, not just the answer. The retrieval trace is your debugging surface. The answer is a symptom.
Knowledge check
Check your understanding
Answer this question before you continue.
Build a Tiny Evaluation Set Before You Debug
You cannot debug what you cannot measure. Write 5–10 questions whose answers live in exactly one chunk, plus one or two questions the corpus genuinely cannot answer.
eval_set = [
{"q": "How long do refunds take?", "expected": "refunds-0"},
{"q": "How fast is express shipping?", "expected": "shipping-1"},
{"q": "What does the warranty cover?", "expected": "warranty-0"},
{"q": "Do you ship to Antarctica?", "expected": None}, # unanswerable
]
def score(k=3):
rows = []
for item in eval_set:
hits = [c["chunk_id"] for c, _ in retrieve(item["q"], k)]
hit = item["expected"] in hits if item["expected"] else None
rows.append((item["q"], item["expected"], hits, hit))
return rows
for row in score():
print(row)
Record the expected source chunk for each question. That is your ground truth for retrieval, independent of whatever the model says. Then score two columns:
- Retrieval hit: is
expectedin the retrieved set? - Grounding pass: is the answer supported by the retrieved text?
The unanswerable questions are not filler. A pipeline that invents a shipping policy for Antarctica has a grounding failure, not a retrieval failure. Those are different bugs with different fixes, and this table is the anchor criterion for everything that follows.
Knowledge check
Check your understanding
Answer this question before you continue.
Read the Trace: Retrieval Failure vs Grounding Failure
Here is the split that stops you from blaming the model for a retrieval problem.
| Failure | What the trace shows | Root cause |
|---|---|---|
| Retrieval failure | Correct chunk never appears in top-k | The model was never given the evidence |
| Grounding failure | Correct chunk retrieved, answer ignores or contradicts it | The model added unsupported detail |
| Quiet third case | Chunk retrieved but truncated or split | Evidence arrived incomplete |
Walk one failing question. Ask "How long do refunds take?" and suppose the trace returns shipping-0, shipping-1, warranty-0 — no refunds-0 at all. The model then answers something plausible about shipping timelines. That is not a hallucination problem. That is a retrieval problem wearing a hallucination costume.
The decision rule: read the retrieved set before you read the answer. If the right chunk is missing, no prompt rewrite will save you. If the right chunk is present and the answer still drifts, now you have a grounding problem worth fixing.
Knowledge check
Check your understanding
Answer this question before you continue.
Fix One Thing at a Time
Change one lever, rerun the evaluation set, compare against the baseline table. Two simultaneous changes make the result unreadable.
Retrieval-side levers:
- Chunk size and overlap — changes what a single chunk can contain.
- Top-k — changes how much evidence reaches the prompt.
- Query phrasing — changes what the embedding matches against.
Grounding-side levers:
- Instruct the model to answer only from the provided context.
- Instruct it to say "I don't know" when the context is insufficient.
Two mistakes eat most beginners here.
Common mistake: Raising top-k to fix a grounding failure. It adds more chunks, more noise, and more chances to drift. It does not fix the cause.
Common mistake: Rewriting the prompt to fix a retrieval failure. The model cannot cite evidence it never received. Fix retrieval first.
A Bounded Improvement, Measured
Take the single failing question from the trace and pull the matching lever only. For the refunds question, the correct chunk was missing entirely, so this is retrieval-side. Suppose the refund answer is split across two sentences and top-k was too small to catch the second. Raise k from 3 to 5 and rerun.
Before and after, same question:
BEFORE: [('shipping-0', 0.71), ('shipping-1', 0.68), ('warranty-0', 0.64)]
AFTER: [('refunds-0', 0.79), ('refunds-1', 0.75), ('shipping-0', 0.71), ...]
Now report the change in the table:
| Metric | Before | After |
|---|---|---|
| Retrieval hits | 2 / 3 answerable | 3 / 3 answerable |
| Grounding passes | 2 / 4 | 3 / 4 |
Interpret honestly. If raising k helped the refunds question but pushed an irrelevant chunk into the Antarctica answer, that is a tradeoff, not a clean win. Say so. And note the limits: a corpus this small does not predict behavior at scale, and an evaluation set this size is a sanity check, not a benchmark. The point is not the number. The point is that you can now see the number move.
Knowledge check
Check your understanding
Answer this question before you continue.
One Experiment to Run Next
Keep the pipeline fixed and change only the chunking boundary. Split on paragraphs instead of sentences, rerun the same evaluation set, and predict the outcome before you run it: which questions should improve, which should degrade, and why.
Then compare your prediction to the observed trace. The gap between what you expected and what the chunks actually did is the real lesson — and it turns this script into a reusable test bench for every retrieval and grounding change you make after it.
Inspect the evidence before you judge the answer. That rule survives long after this corpus does.
Knowledge check
Final check
Finish the article by checking the ideas you just learned.
References
Want a more structured LLMOps path?
Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.
Large Language Models Starter Pack
A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.
- 227-page Illustrated PDF edition
- 12 guided LLM engineering chapters
- Visual concept diagrams
- Self-assessment quizzes
- Bonus deep-dive sections
- Prompt design, structured output, context windows & RAG pipelines
- Agents, tool calling, prompt injection, evaluation & application lifecycles
Coming soon


