Calculate an LLM Context Budget Step by Step
Your request worked yesterday. Today the answer stops mid-sentence, or the API returns an error, and you have no idea which of the four things you sent is…

Key topics
Your request worked yesterday. Today the answer stops mid-sentence, or the API returns an error, and you have no idea which of the four things you sent is to blame. So you do the natural thing: you trim a sentence from your instructions and try again. Sometimes it works. You never learn why.
Here is the reframe that fixes this permanently: the context window is not a limit on what you send. It is a limit on what you send plus what the model writes back. Once you accept that, the problem stops being a feeling and becomes arithmetic. By the end of this article you will be able to write down six numbers, subtract, and know exactly how many tokens to cut and from which block.
Why "Too Long" Is the Wrong Diagnosis
The default beginner belief is that the context window caps your input. That belief survives because it usually works: for a short one-shot question where the answer is a sentence or two, output is a rounding error and input is the whole story.
It breaks the moment any of three things happen. The conversation gets long. You retrieve documents. You run an agent loop that thinks before it answers. In all three cases the model needs room to produce a real answer — a paragraph, a code block, a chain of reasoning — and that room comes out of the same fixed pool as everything you sent.
Picture one desk. Four stacks of paper sit on it: your standing instructions, the conversation so far, the documents you pulled in, and the blank space where the model will write its reply. The desk does not grow when you add to a stack. Add paper to one pile and you are taking desk away from another.
That is the whole model. An LLM context budget is the discipline of deciding, in advance, how much desk each stack gets.
Knowledge check
Check your understanding
Answer this question before you continue.
The Six Numbers You Need
Before the arithmetic, name the parts. Each symbol below maps to something you can literally point at in a request.
C — context capacity. The total tokens the model can hold in a single call, input and output combined. This is a property of the specific model and tier you are calling.
O — reserved output. The space you deliberately hold back for the model's answer. This is a decision you make, not a number the model hands you.
I — instructions. Your system prompt, persona, rules, tool descriptions, and formatting requirements.
H — history. Every prior user message and model reply still sitting in the conversation.
R — retrieved evidence. Documents, search results, file contents, or tool output pulled in for this request.
Q — the current question. The message you are sending right now. It is easy to forget because it feels like the request itself rather than one of its parts, but it consumes tokens like everything else.
A quick bridge on tokens: they are the text units a tokenizer produces, and the same sentence can yield different counts on different models. If you want the full picture of how text becomes tokens and why that matters for limits and cost, the token fundamentals are worth reading first. Here we only need the counting habit.
Visually, think of a single horizontal bar of length C. The right end holds a shaded block of size O — the answer's reserved seat. The left region is everything you send, and inside it sit four stacked bands: I, H, R, and Q.
Knowledge check
Check your understanding
Answer this question before you continue.
Deriving the Input Constraint
Let's build the rule instead of memorizing it.
Step 1. Everything in one call must fit inside capacity:
Step 2. You decide the output size in advance, so output becomes your reserve:
Step 3. Substitute and rearrange:
That right-hand side deserves a name: the input allowance. It is what is left of the desk after you put down the blank page.
Step 4. Input is not one blob. It is four blocks you control:
Read it as a sentence: the four things I control must fit inside what remains after I reserve the answer.
Three assumptions are baked in, and each one can fail.
- O is a target, not a guarantee. The model may stop well short of your reserve, or want more than you gave it.
- The provider may impose a separate maximum output cap. Your reserve cannot exceed it, no matter what the arithmetic says.
- C is a snapshot. It belongs to one model at one tier, and providers change limits over time.
And here is the failure that hides behind correct math: set O too low and the answer gets cut off mid-sentence. Every number checks out. The result is still useless. The arithmetic was right; the reservation was wrong.
Knowledge check
Check your understanding
Answer this question before you continue.
Worked Example: C = 8,000, O = 1,000
Time to run it. Suppose you are calling a model with a context capacity of 8,000 tokens, and you want a solid paragraph or two back, so you reserve 1,000.
| Symbol | Meaning | Tokens |
|---|---|---|
| C | Context capacity | 8,000 |
| O | Reserved output | 1,000 |
| I | Instructions | 1,200 |
| H | History | 3,500 |
| R | Retrieved evidence | 2,800 |
| Q | Current question | 0 |
Input allowance:
Actual input:
Compare:
You are over by 500 tokens. That single number is the deficit, and it is the only number that matters for the fix. Everything else is context.
Now look at where the weight sits. History and evidence together consume 6,300 of the 8,000 tokens — roughly 79% of capacity — before a single word of the answer is written. Instructions are the smallest block at 1,200.
Common mistake: The instinct is to trim the instructions, because that is the part you wrote and it feels editable. But I is the smallest block and the most load-bearing: it tells the model what task it is even doing. The deficit lives in H and R.
Knowledge check
Check your understanding
Answer this question before you continue.
Choosing What to Cut
You need to remove 500 tokens. You have three realistic levers.
Lever 1 — trim history. Drop the oldest turns, or replace them with a short summary. Cut H from 3,500 to 3,000 and input becomes 7,000. The request fits exactly.
Lever 2 — trim evidence. Retrieve fewer or shorter passages. Cut R from 2,800 to 2,300 and input becomes 7,000 again — the same fit, reached from the other side.
Lever 3 — split the difference. Take 250 from each and land at 7,000.
All three pass the arithmetic. Only one is probably right, and the deciding question is this: which block is least likely to change the answer? Recent turns usually carry more signal than old ones. Retrieved passages usually carry more signal than the ones ranked lowest. Cut from the bottom of the value stack, not the top.
Warning: Fitting at exactly 7,000 leaves zero headroom. If the model wants to write 1,100 tokens instead of 1,000, it gets truncated. Aim for a margin — land at 6,800 and keep 200 tokens of slack.
One more thing worth saying plainly: a request that fits is not automatically a good request. Packing the window to the brim with marginal material can degrade the answer even when the arithmetic passes. The goal is the smallest set of high-signal tokens, not the largest set that technically fits.
Where the Numbers Come From — and Why They Move
Treat this calculation as a planning model, not a contract.
Token counts depend on the tokenizer. The same sentence can produce different counts on different models, so a count from one tool is an estimate for another. Context limits and maximum output caps are set by the provider and shift between models, tiers, and over time. Any specific C you write down is a snapshot, not a constant.
The practical approach: count with the tokenizer for the model you are actually calling when you can, and keep a rough estimate in mind when you cannot. The value of the calculation is not precision to the token — it is knowing which block to shrink and by how much before you send the request.
If you are building this into a system, read the usage numbers the provider returns after each call and compare them to your estimate. That gap is your calibration, and it is the only way to make the model of the model more accurate over time.
Reuse the Budget as a Checklist
Here is the whole method, compressed into something you can run on the next request without rereading anything.
- Write down C for the model you are calling.
- Choose O before you write anything else, and choose it generously enough for the answer you actually want.
- Estimate I, H, R, and Q separately rather than guessing at one total.
- Subtract: allowance = C − O. Compare against I + H + R + Q.
- If you are over, cut from the block with the lowest signal value, and leave a margin rather than landing exactly on the limit.
The mistakes worth avoiding are consistent: forgetting that output shares the window, forgetting that your own current question counts as input, reserving an output budget too small to finish the answer, and trimming the instructions the model needs to follow the task at all.
That is the payoff. The reader who can write six numbers and subtract no longer guesses at what to cut — they know the deficit and they know where it lives.
So take a real request that failed. Estimate C, O, I, H, R, and Q, and find your deficit. Then follow the same idea forward: summarization, retrieval tuning, and sliding-window strategies are all ways of keeping H and R from growing without bound in the first place.
Knowledge check
Final check
Finish the article by checking the ideas you just learned.
References
Want a more structured LLMOps path?
Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.
Large Language Models Starter Pack
A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.
- 227-page Illustrated PDF edition
- 12 guided LLM engineering chapters
- Visual concept diagrams
- Self-assessment quizzes
- Bonus deep-dive sections
- Prompt design, structured output, context windows & RAG pipelines
- Agents, tool calling, prompt injection, evaluation & application lifecycles
Coming soon


