Skip to content
intermediate

RAG Chunking Math: Calculate Chunk Counts, Overlap, and Token Use

A 1,500-token document with 512-token chunks and 15% overlap quietly becomes four chunks, and the same paragraph now lives in two of them. Nobody chose…

Published 2026-10-03Updated 2026-10-0410 min read
Technician working on humanoid robot at a tech exhibition in Guimaraes, Portugal.
Technician working on humanoid robot at a tech exhibition in Guimaraes, Portugal. Photo by Rui Dias on Pexels.

A 1,500-token document with 512-token chunks and 15% overlap quietly becomes four chunks, and the same paragraph now lives in two of them. Nobody chose that. The arithmetic chose it.

Most beginners tune chunk size by feel. They try 512, then 256, then 1,024, and judge the results by whether answers "feel better." That is backwards. Before you argue about retrieval quality, compute the three numbers chunk size and overlap already decided for you: how many chunks exist, how many tokens you duplicated, and how much retrieval budget you just spent.

This article assumes you have already decided to use fixed-size chunking. If you are still choosing between fixed-size, recursive, and semantic strategies, that decision comes first. Here, the strategy is fixed and the question is what the numbers do.

The Three Numbers Chunk Size Already Decided

Fixed-size chunking is a sliding window over a token stream. You take the first window, step forward, take the next window, and repeat until the document ends. Everything below follows from five variables:

SymbolMeaning
DDocument length in tokens
CChunk size in tokens
OOverlap in tokens
SStep size, where S = C − O
NChunk count

Two assumptions matter. First, we measure in tokens, not characters or words, because that is what the embedding model and the context window actually consume. Second, we assume the splitter does not snap to sentence or paragraph boundaries. Real splitters often do, so treat the formula as an upper-bound model, not a guarantee.

From these variables you get three outputs: the chunk count N, the duplicated overlap tokens, and the retrieval token budget. Compute them in code, not by asking a model to reason about them. This is deterministic arithmetic, and deterministic arithmetic belongs in a function.

Deriving the Chunk Count Formula

Start with the first chunk. It consumes C tokens and always exists as long as D > 0. That gives us one chunk before any stepping happens.

Each subsequent chunk advances by the step S = C − O. After the first chunk, the remaining content is D − C tokens. The number of additional chunks is however many steps of size S fit into that remainder:

additional chunks=⌈D−CS⌉\text{additional chunks} = \left\lceil \frac{D - C}{S} \right\rceil

Add the first chunk back:

N=1+⌈D−CC−O⌉N = 1 + \left\lceil \frac{D - C}{C - O} \right\rceil

The ceiling appears because the final partial window still becomes a chunk. If 100 tokens remain and the step is 436, you do not get a fraction of a chunk. You get one more chunk holding 100 tokens.

Two degenerate cases are worth naming. If D ≤ C, the document fits in a single chunk and N = 1. If O ≥ C, then S ≤ 0 and the loop never terminates — each step moves backward or nowhere. Overlap must be strictly less than chunk size. A splitter that accepts O = C without complaint is a splitter that will hang or silently truncate.

Read the formula as a tradeoff, not a score. N grows roughly linearly with D. It shrinks as overlap increases, because overlap buys boundary safety by paying in extra chunks.

Knowledge check

Check your understanding

Answer this question before you continue.

A document has 900 tokens, with chunks of 400 tokens and 100 tokens of overlap. How many chunks does the formula predict?
Single Choice

Focus: Calculate a chunk count from document length, chunk size, and overlap using the ceiling rule.

Worked Example: A 1,500-Token Document

A token-axis timeline shows four chunk windows across a 1,500-token document. The first three span 512 tokens; the fourth ends at token 1,500 and is partial. Adjacent windows overlap by 76 tokens and start 436 tokens apart.
The shared axis makes the overlap and short final window visible, clarifying why four 512-token-capacity chunks store 1,728 tokens rather than 2,048.

Take D = 1,500, C = 512, and O = 76 — roughly 15% overlap. Then S = 512 − 76 = 436.

N=1+⌈1500−512436⌉=1+⌈2.27⌉=4N = 1 + \left\lceil \frac{1500 - 512}{436} \right\rceil = 1 + \lceil 2.27 \rceil = 4

Four chunks. Here is where each window starts and ends:

ChunkStart tokenEnd tokenLength
10512512
2436948512
38721,384512
41,3081,500 (partial)192

Now vary one parameter at a time. Drop overlap to zero: S = 512, and N = 1 + ⌈988 / 512⌉ = 3 chunks. Raise overlap to 150: S = 362, and N = 1 + ⌈988 / 362⌉ = 4 chunks. Halve the chunk size to 256 with O = 38: S = 218, and N = 1 + ⌈1244 / 218⌉ = 7 chunks.

Notice what happened. Cutting chunk size in half more than doubled the chunk count. Raising overlap from 76 to 150 bought nothing in chunk count for this document — the ceiling absorbed it. That is the kind of thing you only see when you run the numbers.

The boundary effect is the reason overlap exists. A sentence straddling token 512 gets cut in half by chunk 1 and chunk 2. Overlap means the sentence also appears whole inside chunk 2's window. That is the mechanism, and it is worth remembering that overlap softens boundaries — it does not eliminate them.

Knowledge check

Check your understanding

Answer this question before you continue.

In the worked example, a sentence crosses the end of chunk 1 at token 512 and fits within chunk 2. What does the 76-token overlap do for that sentence?
Scenario Interpretation

Focus: Explain how overlap can preserve a passage that crosses a chunk boundary without eliminating all boundaries.

Counting Duplicated Tokens and the Overlap Tax

Here is where the arithmetic gets subtle, and where a lot of tutorials quietly lie to you. The nominal capacity of your index is N × C. But the actual stored tokens are the sum of the real chunk lengths — and the last chunk is often partial. For our example:

  • Nominal capacity: 4 × 512 = 2,048 tokens
  • Actual stored tokens: 512 + 512 + 512 + 192 = 1,728 tokens
  • Duplicated tokens: 1,728 − 1,500 = 228 tokens

That is about 15% overhead, not the 37% you would get from the naive N × C estimate. The difference matters because it is the difference between a cost model you can trust and one that overstates your bill by more than double.

The general formula for actual duplication is the sum of overlaps between consecutive chunks. When every chunk is full, that sum is (N − 1) × O. When the last chunk is partial, the final overlap is truncated to whatever the partial chunk actually repeats:

duplicated tokens=∑i=1N−1min⁡(O, length of chunk i+1)\text{duplicated tokens} = \sum_{i=1}^{N-1} \min(O,\ \text{length of chunk } i+1)

For our example, the first three overlaps are each 76 tokens (228 total), and the fourth chunk is too short to contribute a full overlap. So duplication lands at 228.

Where does that overhead land? Each duplicated span gets embedded again. Overlap does not just inflate prompt size — it increases embedding calls, vector count, and storage. You pay at index time and again at retrieval time.

Common mistake: Reporting N × C as your stored-token total. That number is the ceiling your index could hold, not what it holds. If your last chunk is partial, N × C overstates both storage and duplication. Compute the actual sum of chunk lengths.

The tradeoff is plain. Overlap protects boundary evidence. The price is paid in duplicate vectors and a larger index. For most documents, 10–15% overlap is enough to soften boundaries without turning your index into a mirror.

Knowledge check

Check your understanding

Answer this question before you continue.

For the worked example, the four actual chunk lengths total 1,728 tokens and the source document has 1,500 tokens. How many tokens are duplicated?
Output Prediction

Focus: Distinguish actual stored-token duplication from nominal chunk capacity when the final chunk is partial.

Turning Chunk Count Into a Retrieval Token Budget

Chunk arithmetic only matters if the retrieved chunks fit in the context window. The budget inequality:

k×C≤context window−system prompt−question−reserved answer spacek \times C \leq \text{context window} - \text{system prompt} - \text{question} - \text{reserved answer space}

Worked check: k = 5, C = 512 gives 2,560 retrieval tokens. Against an 8,192-token window with a 500-token system prompt and a 100-token question, 7,592 tokens remain. It fits, with room for the answer.

Now the failure case. Raise C to 1,024 and k to 8. That demands 8,192 retrieval tokens — the entire window, before the system prompt or the question. It overflows.

Overlap does not change the per-chunk retrieval cost. It changes how many chunks exist and how likely a boundary-spanning passage is retrieved twice. That second effect is subtle: two near-identical chunks can both rank highly for the same query, and you spend two slots on one piece of evidence.

Note: Stuffing more retrieved context into the prompt does not keep improving answers. Beyond a certain point, generation quality can degrade as irrelevant context dilutes the signal. Treat the budget as a constraint to respect, not a target to maximize.

My decision rule: compute the budget before raising k, and prefer raising k over raising C when precision matters. More small chunks give the retriever finer granularity. Fewer large chunks give it more context per slot but fuzzier matches.

Knowledge check

Check your understanding

Answer this question before you continue.

If retrieval returns 5 chunks of 512 tokens each, how many retrieval tokens must the context budget accommodate?
Output Prediction

Focus: Calculate the retrieval-token demand from the number and size of retrieved chunks.

What the Math Cannot Tell You

The formula predicts how many chunks exist and what they cost. It says nothing about whether the right chunk ranks highly for a given query.

This is the boundary that trips people up. Chunk size effects on retrieval quality can be small or even negligible when the retrieval goal is document-level rather than passage-level. If any chunk from the correct document counts as a success, larger chunks still contain the relevant content, just with more surrounding context. The metric you choose changes the conclusion.

Boundary effects are semantic, not arithmetic. A chunk can be exactly 512 tokens and still cut a definition away from the term it defines. The math says the chunk is well-sized. The reader says the answer is wrong.

And embedding similarity is not factual correctness. A well-sized chunk can be retrieved for the wrong reason — vector closeness that looks right but misses the point.

State it plainly: the count and budget math is deterministic. The relevance consequence is empirical and must be measured on your own documents. Do the math to bound cost and catch overflow, then evaluate retrieval separately.

A Repeatable Procedure for Your Own Documents

Turn the derivation into a checklist you can run before tuning anything:

  1. Measure D in tokens with the same tokenizer the embedding model uses, not by word count.
  2. Pick C and O, compute S = C − O, then compute N. Reject any configuration where O ≥ C.
  3. Sum the actual chunk lengths to get stored tokens, then subtract D to get real duplication. Do not use N × C as your stored-token total.
  4. Check k × C against the available context window before running retrieval.
  5. Only after those numbers pass, evaluate retrieval quality and adjust one parameter at a time.

Three mistakes to avoid: mixing character counts with token counts, treating overlap as free, and tuning chunk size against a metric that does not match your retrieval goal.

You now own two separate jobs. The arithmetic is deterministic and belongs in code — write a function that takes D, C, and O and returns N, actual stored tokens, real duplication, and the budget check. The retrieval evaluation is empirical and cannot be inferred from chunk counts. Compute the numbers for one real document today, then measure retrieval quality as the next step. The math tells you what you can afford. Only measurement tells you whether it worked.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

A chunking calculation shows that a configuration fits the context budget. What can you conclude about whether it will retrieve relevant evidence?
Question 1 of 2Misconception Check

Focus: Separate deterministic chunk-count and cost calculations from empirical retrieval-quality conclusions.

Which workflow best follows the article's procedure before tuning a chunking configuration?
Question 2 of 2Comparison Reasoning

Focus: Choose a workflow that follows the article's procedure for calculating chunk mechanics before assessing retrieval quality.

References

  1. RAG chunking explained: chunk size, overlap, and what it costsflaviocopes.com
  2. Evaluate Your Own RAG: Why Best Practices Failed Ushuggingface.co
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.