Skip to content
intermediate

Test RAG Chunking Strategies on a Small Document Set

Retrieval misses rarely announce their cause. The answer was in the document, the query was reasonable, and the model still returned something useless. The…

Published 2026-10-03Updated 2026-10-0411 min read
A beautiful sunset with vibrant colors over a tranquil body of water.
A beautiful sunset with vibrant colors over a tranquil body of water. Photo by Stephen Andrews on Pexels.

Retrieval misses rarely announce their cause. The answer was in the document, the query was reasonable, and the model still returned something useless. The reflex is to blame the embedding model, the vector store, or the LLM. Most of the time, the problem is upstream: the answer was cut in half by a chunk boundary, or the chunk that held it ranked below the cutoff.

Chunking is a preprocessing decision, which means it is measurable. You can test it this afternoon with no API keys, no vector database, and no embeddings. What follows is a small controlled experiment: a fixed corpus, a fixed query set, a few chunk-size and overlap settings, and a scorer transparent enough that you can see exactly why every passage ranked where it did.

This tells you what your corpus and your queries do. It does not tell you the best chunk size in general. Nothing does.

What This Experiment Can and Cannot Tell You

Before writing a line of code, set the boundary. This is a bounded comparison on one small fixture, not a benchmark.

The scorer here is token overlap: it counts how many query words appear in a chunk. That makes it deterministic, dependency-free, and fully inspectable. It also means it favors passages that share vocabulary with the query and will miss paraphrases that an embedding model would catch. That is a known limitation, not a bug. You are testing chunk boundaries, not semantic understanding.

What this design deliberately excludes: embedding quality, reranking, generation quality, and cost. Each is a separate variable with its own failure modes. If you change all of them at once, you learn nothing about any of them.

If you already know the chunking tradeoffs — smaller chunks sharpen matching but fragment context, larger chunks preserve context but dilute the score — this is where you stop reasoning about them in the abstract and start measuring them on your own material.

Set Up the Corpus, Queries, and Relevance Labels

Environment: Python 3.11, standard library only. No pip installs, no API keys, no external services.

The corpus is a handful of short plain-text documents — small enough to read end to end and reason about by hand. Keep them short. If you cannot hold the whole corpus in your head, you cannot tell whether a retrieval miss came from chunking or from your own confusion about what the documents say.

The query fixture is a small set of realistic queries. Each query gets a relevance label written before you run anything: which document, and ideally which span, should be retrieved.

Warning: Labels written after you see results turn an experiment into a rationalization. Write them first. If a label turns out to be wrong, note it and restart — do not quietly edit it to match what the retriever returned.

Freeze the corpus and queries for the whole run. Changing inputs mid-experiment invalidates every comparison you already made.

CORPUS = {
    "billing.txt": (
        "Invoices are generated on the first business day of each month. "
        "Payment terms are net 30 from the invoice date. "
        "Late payments incur a 1.5 percent monthly fee. "
        "Refunds are processed within ten business days of approval."
    ),
    "shipping.txt": (
        "Standard shipping arrives in three to five business days. "
        "Expedited shipping arrives in one to two business days. "
        "Orders over fifty dollars ship free within the continental region. "
        "International orders may face customs delays."
    ),
    "returns.txt": (
        "Returns are accepted within thirty days of delivery. "
        "Items must be unused and in original packaging. "
        "Return shipping is free for defective items. "
        "Store credit is issued for non-defective returns."
    ),
}

QUERIES = [
    {"query": "how long do refunds take", "label_doc": "billing.txt", "label_span": "ten business days"},
    {"query": "when is expedited shipping", "label_doc": "shipping.txt", "label_span": "one to two business days"},
    {"query": "return shipping cost for defective items", "label_doc": "returns.txt", "label_span": "free for defective items"},
    {"query": "late payment fee", "label_doc": "billing.txt", "label_span": "1.5 percent monthly fee"},
]

Four queries is enough to expose a boundary failure. It is not enough to declare a winner.

Build the Chunker: Size and Overlap as Two Knobs

The chunker is one function with two parameters. Keep it that small.

def chunk_text(text, chunk_size, overlap):
    stride = chunk_size - overlap
    if stride <= 0:
        raise ValueError("overlap must be smaller than chunk_size")
    chunks = []
    start = 0
    while start < len(text):
        end = start + chunk_size
        chunks.append({"text": text[start:end], "start": start, "end": end})
        start += stride
    return chunks

The arithmetic that matters is stride = chunk_size - overlap. If overlap equals or exceeds chunk size, the loop never advances and you get an infinite loop or a single chunk. Guard it explicitly.

Keep the start and end offsets on every chunk. Without them, you cannot tell whether a miss came from bad ranking or from a boundary that cut the answer in half. Offsets are how you turn a vague failure into a diagnosable one.

The strategy set to compare: a few chunk sizes crossed with a couple of overlap values. Print chunk counts and chunk lengths per configuration before any retrieval happens — if a configuration produces one giant chunk or fifty tiny ones, you want to see that before you start interpreting scores.

CONFIGS = [
    {"name": "size=80 overlap=0",  "chunk_size": 80,  "overlap": 0},
    {"name": "size=80 overlap=20", "chunk_size": 80,  "overlap": 20},
    {"name": "size=160 overlap=0", "chunk_size": 160, "overlap": 0},
    {"name": "size=160 overlap=40","chunk_size": 160, "overlap": 40},
]

Knowledge check

Check your understanding

Answer this question before you continue.

A run is configured with `chunk_size=80` and `overlap=80`. What should the chunker do with this configuration?
Debugging

Focus: Identify why a chunking configuration with overlap equal to chunk size cannot make forward progress.

Score Retrieval with a Transparent Token-Overlap Scorer

Normalize the text, tokenize on whitespace, and score each chunk by how much of the query it contains.

import re

def tokenize(text):
    return re.findall(r"[a-z0-9]+", text.lower())

def overlap_score(query, chunk_text):
    q_tokens = set(tokenize(query))
    c_tokens = set(tokenize(chunk_text))
    if not q_tokens:
        return 0.0
    return len(q_tokens & c_tokens) / len(q_tokens)

The score is the fraction of query tokens present in the chunk. A score of 1.0 means every query word appears somewhere in the chunk. It says nothing about whether those words are in the right order, in the right sentence, or answering the question. That is the point — the number is honest about how little it knows.

Rank chunks per query, take the top-k, and print the score, offsets, and a short excerpt for each returned passage.

def retrieve(query, chunks, k=2):
    scored = [(overlap_score(query, c["text"]), c) for c in chunks]
    scored.sort(key=lambda pair: pair[0], reverse=True)
    return scored[:k]

Common mistake: A chunk can score high on shared stopwords or boilerplate while containing none of the answer. If your query is "how long do refunds take," a chunk full of "do" and "take" can outrank the sentence that actually says "ten business days." This is exactly why the next section checks evidence, not just rank.

Knowledge check

Check your understanding

Answer this question before you continue.

Using the article's tokenizer and scorer, what score does this chunk receive for the query?
Output Prediction

Focus: Calculate the scorer's query-token coverage for a query and a candidate chunk.

query: how long do refunds take
chunk: Refunds take ten business days.

Assemble the Runner and Evaluate Against Labels

A flow runs from corpus through chunk configurations to ranked top-k results, then branches into two checks: whether the labeled document appears and whether the labeled span is intact in one returned chunk.
Evaluate document presence and intact answer context separately to tell ranking misses from chunk-boundary failures.

The helpers above are not yet an experiment. You need one loop that chunks every document while keeping its source identity, retrieves for each query, and checks the returned passages against the labels. Without source-document metadata, a returned chunk cannot be matched to its label at all.

def build_chunks(corpus, chunk_size, overlap):
    all_chunks = []
    for doc_name, text in corpus.items():
        for chunk in chunk_text(text, chunk_size, overlap):
            chunk["doc"] = doc_name
            all_chunks.append(chunk)
    return all_chunks

def span_present(chunk, label_span):
    return label_span.lower() in chunk["text"].lower()

def evaluate(config):
    chunks = build_chunks(CORPUS, config["chunk_size"], config["overlap"])
    evidence_hits = 0
    context_hits = 0
    for item in QUERIES:
        results = retrieve(item["query"], chunks, k=2)
        top_docs = [c["doc"] for _, c in results]
        found = item["label_doc"] in top_docs
        complete = any(
            c["doc"] == item["label_doc"] and span_present(c, item["label_span"])
            for _, c in results
        )
        evidence_hits += int(found)
        context_hits += int(complete)
        print(f"  {item['query']!r} -> evidence={found} complete={complete}")
    print(f"{config['name']}: evidence {evidence_hits}/{len(QUERIES)}, "
          f"context {context_hits}/{len(QUERIES)}")

for config in CONFIGS:
    evaluate(config)

Run it. The output is the experiment. Each line tells you whether the labeled document appeared in the top-k and whether the labeled span survived intact inside a returned chunk. The two counts at the bottom are the numbers that go into your table.

For each query, ask two separate questions. They have different failure signatures, and conflating them is how people end up tuning the wrong knob.

Evidence relevance: did the top-k set include the labeled document at all? A miss here is a recall problem — the right text exists in some chunk, but that chunk did not rank high enough.

Context completeness: when the evidence is present, is the labeled span intact inside a single returned chunk? A miss here is a boundary problem — the answer is split across two chunks, or buried under so much unrelated text that a downstream model would struggle to use it.

The two failures look different in the output:

  • A boundary cut shows up as partial evidence in one chunk plus a neighboring chunk holding the rest. The label span straddles an offset.
  • A ranking miss shows up as the correct chunk existing in the index but scoring below the top-k cutoff. You can see it by printing the full ranked list, not just the winners.

Complete the results table from your own run. One row per configuration, with a short note on what went wrong when something did.

ConfigurationEvidence foundContext completeNote
size=80 overlap=0
size=80 overlap=20
size=160 overlap=0
size=160 overlap=40

Fill every cell from the printed output. If a configuration ties on both counts, note that too — a tie is a result, not a gap.

Knowledge check

Check your understanding

Answer this question before you continue.

For a query, a top-k result comes from the labeled document, but none of the returned chunks contains the complete labeled span. How should the runner record the result?
Scenario Interpretation

Focus: Distinguish retrieval of a labeled document from intact retrieval of its labeled answer span.

Read the Table Without Overclaiming

Expect small differences. Published RAG evaluations have found chunk size having limited impact when the retrieval goal is document-level rather than exact-span, and a four-query fixture will behave the same way — noisily.

Separate what your run shows from what it cannot. One corpus, one scorer, one query set proves nothing about other documents or embedding-based retrieval. What it does show is the shape of the tradeoff: smaller chunks tend to sharpen matching but fragment context; larger chunks preserve context but dilute the score.

The decision rule: pick the configuration that keeps labeled evidence intact across your queries, then stop tuning. Move on to the next bottleneck — embeddings, reranking, or generation — because those are where the remaining error usually lives.

Common mistake: Chasing a two-point difference on a ten-query fixture is overfitting to noise, not optimization. If two configurations tie on evidence and context, pick the simpler one and walk away.

Knowledge check

Check your understanding

Answer this question before you continue.

Two configurations tie on both evidence-found and context-complete counts in this small fixture. What is the article's recommended response?
Comparison Reasoning

Focus: Apply the article's decision rule when configurations tie on evidence and context results.

Extend the Experiment

One meaningful modification beats five decorative ones. Pick the hypothesis that matches your real pipeline.

Add a structure-aware splitter. Break on paragraph or heading boundaries and compare it against the fixed-size configurations on the same queries. This is the contrast that tells you whether your documents have natural seams worth respecting.

Add a query that spans two sections. Watch how each strategy handles evidence that legitimately needs more than one chunk. This is where fixed-size chunking tends to fail loudly, and where overlap either saves you or just adds redundancy.

Swap the scorer for a stemming or stopword filter. Observe how much of the ranking was driven by common words. If removing stopwords reshuffles the top-k, your scorer was measuring noise.

Log the failure cases instead of deleting them. The queries that break are the ones that tell you what your real pipeline will need. A miss is data. A deleted miss is a lie you tell yourself about your retrieval quality.

Where This Leaves You

You now have a reusable harness: fixed corpus, fixed queries, labeled evidence, transparent scorer, and a runner that prints the two numbers that matter. Point it at your own documents before you touch embeddings or reranking. Chunking is the cheapest variable to test and the easiest to blame wrongly.

The rule holds: keep evidence intact, then stop tuning. Run the same harness against a real document set next — a folder of your own notes, docs, or support tickets — and let the table tell you which configuration survives contact with your actual material.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

After seeing retrieval results, you suspect one prewritten relevance label is wrong. What should you do to preserve the experiment's integrity?
Question 1 of 2Misconception Check

Focus: Preserve a controlled evaluation by handling a relevance label that appears incorrect after a run.

Your documents have meaningful headings and paragraphs, and you want to test whether those natural seams matter for retrieval. Which modification best tests that hypothesis?
Question 2 of 2Scenario Interpretation

Focus: Choose a controlled experiment modification for testing whether natural document boundaries improve chunking.

References

  1. Evaluate Your Own RAG: Why Best Practices Failed Ushuggingface.co
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.