Test RAG Chunking Strategies on a Small Document Set
Retrieval misses rarely announce their cause. The answer was in the document, the query was reasonable, and the model still returned something useless. The…

Key topics
Retrieval misses rarely announce their cause. The answer was in the document, the query was reasonable, and the model still returned something useless. The reflex is to blame the embedding model, the vector store, or the LLM. Most of the time, the problem is upstream: the answer was cut in half by a chunk boundary, or the chunk that held it ranked below the cutoff.
Chunking is a preprocessing decision, which means it is measurable. You can test it this afternoon with no API keys, no vector database, and no embeddings. What follows is a small controlled experiment: a fixed corpus, a fixed query set, a few chunk-size and overlap settings, and a scorer transparent enough that you can see exactly why every passage ranked where it did.
This tells you what your corpus and your queries do. It does not tell you the best chunk size in general. Nothing does.
What This Experiment Can and Cannot Tell You
Before writing a line of code, set the boundary. This is a bounded comparison on one small fixture, not a benchmark.
The scorer here is token overlap: it counts how many query words appear in a chunk. That makes it deterministic, dependency-free, and fully inspectable. It also means it favors passages that share vocabulary with the query and will miss paraphrases that an embedding model would catch. That is a known limitation, not a bug. You are testing chunk boundaries, not semantic understanding.
What this design deliberately excludes: embedding quality, reranking, generation quality, and cost. Each is a separate variable with its own failure modes. If you change all of them at once, you learn nothing about any of them.
If you already know the chunking tradeoffs — smaller chunks sharpen matching but fragment context, larger chunks preserve context but dilute the score — this is where you stop reasoning about them in the abstract and start measuring them on your own material.
Set Up the Corpus, Queries, and Relevance Labels
Environment: Python 3.11, standard library only. No pip installs, no API keys, no external services.
The corpus is a handful of short plain-text documents — small enough to read end to end and reason about by hand. Keep them short. If you cannot hold the whole corpus in your head, you cannot tell whether a retrieval miss came from chunking or from your own confusion about what the documents say.
The query fixture is a small set of realistic queries. Each query gets a relevance label written before you run anything: which document, and ideally which span, should be retrieved.
Warning: Labels written after you see results turn an experiment into a rationalization. Write them first. If a label turns out to be wrong, note it and restart — do not quietly edit it to match what the retriever returned.
Freeze the corpus and queries for the whole run. Changing inputs mid-experiment invalidates every comparison you already made.
CORPUS = {
"billing.txt": (
"Invoices are generated on the first business day of each month. "
"Payment terms are net 30 from the invoice date. "
"Late payments incur a 1.5 percent monthly fee. "
"Refunds are processed within ten business days of approval."
),
"shipping.txt": (
"Standard shipping arrives in three to five business days. "
"Expedited shipping arrives in one to two business days. "
"Orders over fifty dollars ship free within the continental region. "
"International orders may face customs delays."
),
"returns.txt": (
"Returns are accepted within thirty days of delivery. "
"Items must be unused and in original packaging. "
"Return shipping is free for defective items. "
"Store credit is issued for non-defective returns."
),
}
QUERIES = [
{"query": "how long do refunds take", "label_doc": "billing.txt", "label_span": "ten business days"},
{"query": "when is expedited shipping", "label_doc": "shipping.txt", "label_span": "one to two business days"},
{"query": "return shipping cost for defective items", "label_doc": "returns.txt", "label_span": "free for defective items"},
{"query": "late payment fee", "label_doc": "billing.txt", "label_span": "1.5 percent monthly fee"},
]
Four queries is enough to expose a boundary failure. It is not enough to declare a winner.
Build the Chunker: Size and Overlap as Two Knobs
The chunker is one function with two parameters. Keep it that small.
def chunk_text(text, chunk_size, overlap):
stride = chunk_size - overlap
if stride <= 0:
raise ValueError("overlap must be smaller than chunk_size")
chunks = []
start = 0
while start < len(text):
end = start + chunk_size
chunks.append({"text": text[start:end], "start": start, "end": end})
start += stride
return chunks
The arithmetic that matters is stride = chunk_size - overlap. If overlap equals or exceeds chunk size, the loop never advances and you get an infinite loop or a single chunk. Guard it explicitly.
Keep the start and end offsets on every chunk. Without them, you cannot tell whether a miss came from bad ranking or from a boundary that cut the answer in half. Offsets are how you turn a vague failure into a diagnosable one.
The strategy set to compare: a few chunk sizes crossed with a couple of overlap values. Print chunk counts and chunk lengths per configuration before any retrieval happens — if a configuration produces one giant chunk or fifty tiny ones, you want to see that before you start interpreting scores.
CONFIGS = [
{"name": "size=80 overlap=0", "chunk_size": 80, "overlap": 0},
{"name": "size=80 overlap=20", "chunk_size": 80, "overlap": 20},
{"name": "size=160 overlap=0", "chunk_size": 160, "overlap": 0},
{"name": "size=160 overlap=40","chunk_size": 160, "overlap": 40},
]
Knowledge check
Check your understanding
Answer this question before you continue.
Score Retrieval with a Transparent Token-Overlap Scorer
Normalize the text, tokenize on whitespace, and score each chunk by how much of the query it contains.
import re
def tokenize(text):
return re.findall(r"[a-z0-9]+", text.lower())
def overlap_score(query, chunk_text):
q_tokens = set(tokenize(query))
c_tokens = set(tokenize(chunk_text))
if not q_tokens:
return 0.0
return len(q_tokens & c_tokens) / len(q_tokens)
The score is the fraction of query tokens present in the chunk. A score of 1.0 means every query word appears somewhere in the chunk. It says nothing about whether those words are in the right order, in the right sentence, or answering the question. That is the point — the number is honest about how little it knows.
Rank chunks per query, take the top-k, and print the score, offsets, and a short excerpt for each returned passage.
def retrieve(query, chunks, k=2):
scored = [(overlap_score(query, c["text"]), c) for c in chunks]
scored.sort(key=lambda pair: pair[0], reverse=True)
return scored[:k]
Common mistake: A chunk can score high on shared stopwords or boilerplate while containing none of the answer. If your query is "how long do refunds take," a chunk full of "do" and "take" can outrank the sentence that actually says "ten business days." This is exactly why the next section checks evidence, not just rank.
Knowledge check
Check your understanding
Answer this question before you continue.
Assemble the Runner and Evaluate Against Labels
The helpers above are not yet an experiment. You need one loop that chunks every document while keeping its source identity, retrieves for each query, and checks the returned passages against the labels. Without source-document metadata, a returned chunk cannot be matched to its label at all.
def build_chunks(corpus, chunk_size, overlap):
all_chunks = []
for doc_name, text in corpus.items():
for chunk in chunk_text(text, chunk_size, overlap):
chunk["doc"] = doc_name
all_chunks.append(chunk)
return all_chunks
def span_present(chunk, label_span):
return label_span.lower() in chunk["text"].lower()
def evaluate(config):
chunks = build_chunks(CORPUS, config["chunk_size"], config["overlap"])
evidence_hits = 0
context_hits = 0
for item in QUERIES:
results = retrieve(item["query"], chunks, k=2)
top_docs = [c["doc"] for _, c in results]
found = item["label_doc"] in top_docs
complete = any(
c["doc"] == item["label_doc"] and span_present(c, item["label_span"])
for _, c in results
)
evidence_hits += int(found)
context_hits += int(complete)
print(f" {item['query']!r} -> evidence={found} complete={complete}")
print(f"{config['name']}: evidence {evidence_hits}/{len(QUERIES)}, "
f"context {context_hits}/{len(QUERIES)}")
for config in CONFIGS:
evaluate(config)
Run it. The output is the experiment. Each line tells you whether the labeled document appeared in the top-k and whether the labeled span survived intact inside a returned chunk. The two counts at the bottom are the numbers that go into your table.
For each query, ask two separate questions. They have different failure signatures, and conflating them is how people end up tuning the wrong knob.
Evidence relevance: did the top-k set include the labeled document at all? A miss here is a recall problem — the right text exists in some chunk, but that chunk did not rank high enough.
Context completeness: when the evidence is present, is the labeled span intact inside a single returned chunk? A miss here is a boundary problem — the answer is split across two chunks, or buried under so much unrelated text that a downstream model would struggle to use it.
The two failures look different in the output:
- A boundary cut shows up as partial evidence in one chunk plus a neighboring chunk holding the rest. The label span straddles an offset.
- A ranking miss shows up as the correct chunk existing in the index but scoring below the top-k cutoff. You can see it by printing the full ranked list, not just the winners.
Complete the results table from your own run. One row per configuration, with a short note on what went wrong when something did.
| Configuration | Evidence found | Context complete | Note |
|---|---|---|---|
| size=80 overlap=0 | |||
| size=80 overlap=20 | |||
| size=160 overlap=0 | |||
| size=160 overlap=40 |
Fill every cell from the printed output. If a configuration ties on both counts, note that too — a tie is a result, not a gap.
Knowledge check
Check your understanding
Answer this question before you continue.
Read the Table Without Overclaiming
Expect small differences. Published RAG evaluations have found chunk size having limited impact when the retrieval goal is document-level rather than exact-span, and a four-query fixture will behave the same way — noisily.
Separate what your run shows from what it cannot. One corpus, one scorer, one query set proves nothing about other documents or embedding-based retrieval. What it does show is the shape of the tradeoff: smaller chunks tend to sharpen matching but fragment context; larger chunks preserve context but dilute the score.
The decision rule: pick the configuration that keeps labeled evidence intact across your queries, then stop tuning. Move on to the next bottleneck — embeddings, reranking, or generation — because those are where the remaining error usually lives.
Common mistake: Chasing a two-point difference on a ten-query fixture is overfitting to noise, not optimization. If two configurations tie on evidence and context, pick the simpler one and walk away.
Knowledge check
Check your understanding
Answer this question before you continue.
Extend the Experiment
One meaningful modification beats five decorative ones. Pick the hypothesis that matches your real pipeline.
Add a structure-aware splitter. Break on paragraph or heading boundaries and compare it against the fixed-size configurations on the same queries. This is the contrast that tells you whether your documents have natural seams worth respecting.
Add a query that spans two sections. Watch how each strategy handles evidence that legitimately needs more than one chunk. This is where fixed-size chunking tends to fail loudly, and where overlap either saves you or just adds redundancy.
Swap the scorer for a stemming or stopword filter. Observe how much of the ranking was driven by common words. If removing stopwords reshuffles the top-k, your scorer was measuring noise.
Log the failure cases instead of deleting them. The queries that break are the ones that tell you what your real pipeline will need. A miss is data. A deleted miss is a lie you tell yourself about your retrieval quality.
Where This Leaves You
You now have a reusable harness: fixed corpus, fixed queries, labeled evidence, transparent scorer, and a runner that prints the two numbers that matter. Point it at your own documents before you touch embeddings or reranking. Chunking is the cheapest variable to test and the easiest to blame wrongly.
The rule holds: keep evidence intact, then stop tuning. Run the same harness against a real document set next — a folder of your own notes, docs, or support tickets — and let the table tell you which configuration survives contact with your actual material.
Knowledge check
Final check
Finish the article by checking the ideas you just learned.
References
Want a more structured LLMOps path?
Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.
Large Language Models Starter Pack
A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.
- 227-page Illustrated PDF edition
- 12 guided LLM engineering chapters
- Visual concept diagrams
- Self-assessment quizzes
- Bonus deep-dive sections
- Prompt design, structured output, context windows & RAG pipelines
- Agents, tool calling, prompt injection, evaluation & application lifecycles
Coming soon


