Skip to content
beginner

Calculate an LLM Context Budget Step by Step

Your request worked yesterday. Today the answer stops mid-sentence, or the API returns an error, and you have no idea which of the four things you sent is…

Published 2026-10-03Updated 2026-10-049 min read
Detailed image of a computer keyboard with blue LED backlighting, highlighting keys.
Detailed image of a computer keyboard with blue LED backlighting, highlighting keys. Photo by Marta Branco on Pexels.

Your request worked yesterday. Today the answer stops mid-sentence, or the API returns an error, and you have no idea which of the four things you sent is to blame. So you do the natural thing: you trim a sentence from your instructions and try again. Sometimes it works. You never learn why.

Here is the reframe that fixes this permanently: the context window is not a limit on what you send. It is a limit on what you send plus what the model writes back. Once you accept that, the problem stops being a feeling and becomes arithmetic. By the end of this article you will be able to write down six numbers, subtract, and know exactly how many tokens to cut and from which block.

Why "Too Long" Is the Wrong Diagnosis

The default beginner belief is that the context window caps your input. That belief survives because it usually works: for a short one-shot question where the answer is a sentence or two, output is a rounding error and input is the whole story.

It breaks the moment any of three things happen. The conversation gets long. You retrieve documents. You run an agent loop that thinks before it answers. In all three cases the model needs room to produce a real answer — a paragraph, a code block, a chain of reasoning — and that room comes out of the same fixed pool as everything you sent.

Picture one desk. Four stacks of paper sit on it: your standing instructions, the conversation so far, the documents you pulled in, and the blank space where the model will write its reply. The desk does not grow when you add to a stack. Add paper to one pile and you are taking desk away from another.

That is the whole model. An LLM context budget is the discipline of deciding, in advance, how much desk each stack gets.

Knowledge check

Check your understanding

Answer this question before you continue.

A request fits when you count only the text sent, but the answer is cut off. Which explanation best matches the article’s context-budget model?
Misconception Check

Focus: Recognize why a context limit applies to the combined input and generated output.

The Six Numbers You Need

Before the arithmetic, name the parts. Each symbol below maps to something you can literally point at in a request.

C — context capacity. The total tokens the model can hold in a single call, input and output combined. This is a property of the specific model and tier you are calling.

O — reserved output. The space you deliberately hold back for the model's answer. This is a decision you make, not a number the model hands you.

I — instructions. Your system prompt, persona, rules, tool descriptions, and formatting requirements.

H — history. Every prior user message and model reply still sitting in the conversation.

R — retrieved evidence. Documents, search results, file contents, or tool output pulled in for this request.

Q — the current question. The message you are sending right now. It is easy to forget because it feels like the request itself rather than one of its parts, but it consumes tokens like everything else.

A quick bridge on tokens: they are the text units a tokenizer produces, and the same sentence can yield different counts on different models. If you want the full picture of how text becomes tokens and why that matters for limits and cost, the token fundamentals are worth reading first. Here we only need the counting habit.

Visually, think of a single horizontal bar of length C. The right end holds a shaded block of size O — the answer's reserved seat. The left region is everything you send, and inside it sit four stacked bands: I, H, R, and Q.

Knowledge check

Check your understanding

Answer this question before you continue.

When estimating the input for a request, which component is easy to overlook because it feels like the request itself?
Single Choice

Focus: Identify the current question as a token-consuming input component in a budget estimate.

Deriving the Input Constraint

Let's build the rule instead of memorizing it.

Step 1. Everything in one call must fit inside capacity:

input+output≤C\text{input} + \text{output} \le C

Step 2. You decide the output size in advance, so output becomes your reserve:

output=O\text{output} = O

Step 3. Substitute and rearrange:

input≤C−O\text{input} \le C - O

That right-hand side deserves a name: the input allowance. It is what is left of the desk after you put down the blank page.

Step 4. Input is not one blob. It is four blocks you control:

I+H+R+Q≤C−OI + H + R + Q \le C - O

Read it as a sentence: the four things I control must fit inside what remains after I reserve the answer.

Three assumptions are baked in, and each one can fail.

  • O is a target, not a guarantee. The model may stop well short of your reserve, or want more than you gave it.
  • The provider may impose a separate maximum output cap. Your reserve cannot exceed it, no matter what the arithmetic says.
  • C is a snapshot. It belongs to one model at one tier, and providers change limits over time.

And here is the failure that hides behind correct math: set O too low and the answer gets cut off mid-sentence. Every number checks out. The result is still useless. The arithmetic was right; the reservation was wrong.

Knowledge check

Check your understanding

Answer this question before you continue.

Which expression states the article’s input constraint after reserving O tokens for output?
Single Choice

Focus: Use the derived input constraint with all four input components and an explicit output reserve.

Worked Example: C = 8,000, O = 1,000

A proportional context bar shows 8,000 tokens split into a 7,000-token input allowance and 1,000-token output reserve. A second bar stacks 1,200 instruction, 3,500 history, and 2,800 evidence tokens; it extends 500 tokens past the allowance. The current question is zero.
The request is 500 tokens over budget; history and retrieved evidence account for most of its input.

Time to run it. Suppose you are calling a model with a context capacity of 8,000 tokens, and you want a solid paragraph or two back, so you reserve 1,000.

SymbolMeaningTokens
CContext capacity8,000
OReserved output1,000
IInstructions1,200
HHistory3,500
RRetrieved evidence2,800
QCurrent question0

Input allowance:

C−O=8,000−1,000=7,000C - O = 8{,}000 - 1{,}000 = 7{,}000

Actual input:

I+H+R+Q=1,200+3,500+2,800+0=7,500I + H + R + Q = 1{,}200 + 3{,}500 + 2{,}800 + 0 = 7{,}500

Compare:

7,500>7,0007{,}500 > 7{,}000

You are over by 500 tokens. That single number is the deficit, and it is the only number that matters for the fix. Everything else is context.

Now look at where the weight sits. History and evidence together consume 6,300 of the 8,000 tokens — roughly 79% of capacity — before a single word of the answer is written. Instructions are the smallest block at 1,200.

Common mistake: The instinct is to trim the instructions, because that is the part you wrote and it feels editable. But I is the smallest block and the most load-bearing: it tells the model what task it is even doing. The deficit lives in H and R.

Knowledge check

Check your understanding

Answer this question before you continue.

In the worked example, the input is 7,500 tokens and the input allowance is 7,000. Which change brings the input exactly to the allowance?
Scenario Interpretation

Focus: Calculate the input deficit in the worked numerical example and identify a cut that brings the request to the allowance.

Choosing What to Cut

You need to remove 500 tokens. You have three realistic levers.

Lever 1 — trim history. Drop the oldest turns, or replace them with a short summary. Cut H from 3,500 to 3,000 and input becomes 7,000. The request fits exactly.

Lever 2 — trim evidence. Retrieve fewer or shorter passages. Cut R from 2,800 to 2,300 and input becomes 7,000 again — the same fit, reached from the other side.

Lever 3 — split the difference. Take 250 from each and land at 7,000.

All three pass the arithmetic. Only one is probably right, and the deciding question is this: which block is least likely to change the answer? Recent turns usually carry more signal than old ones. Retrieved passages usually carry more signal than the ones ranked lowest. Cut from the bottom of the value stack, not the top.

Warning: Fitting at exactly 7,000 leaves zero headroom. If the model wants to write 1,100 tokens instead of 1,000, it gets truncated. Aim for a margin — land at 6,800 and keep 200 tokens of slack.

One more thing worth saying plainly: a request that fits is not automatically a good request. Packing the window to the brim with marginal material can degrade the answer even when the arithmetic passes. The goal is the smallest set of high-signal tokens, not the largest set that technically fits.

Where the Numbers Come From — and Why They Move

Treat this calculation as a planning model, not a contract.

Token counts depend on the tokenizer. The same sentence can produce different counts on different models, so a count from one tool is an estimate for another. Context limits and maximum output caps are set by the provider and shift between models, tiers, and over time. Any specific C you write down is a snapshot, not a constant.

The practical approach: count with the tokenizer for the model you are actually calling when you can, and keep a rough estimate in mind when you cannot. The value of the calculation is not precision to the token — it is knowing which block to shrink and by how much before you send the request.

If you are building this into a system, read the usage numbers the provider returns after each call and compare them to your estimate. That gap is your calibration, and it is the only way to make the model of the model more accurate over time.

Reuse the Budget as a Checklist

Here is the whole method, compressed into something you can run on the next request without rereading anything.

  1. Write down C for the model you are calling.
  2. Choose O before you write anything else, and choose it generously enough for the answer you actually want.
  3. Estimate I, H, R, and Q separately rather than guessing at one total.
  4. Subtract: allowance = C − O. Compare against I + H + R + Q.
  5. If you are over, cut from the block with the lowest signal value, and leave a margin rather than landing exactly on the limit.

The mistakes worth avoiding are consistent: forgetting that output shares the window, forgetting that your own current question counts as input, reserving an output budget too small to finish the answer, and trimming the instructions the model needs to follow the task at all.

That is the payoff. The reader who can write six numbers and subtract no longer guesses at what to cut — they know the deficit and they know where it lives.

So take a real request that failed. Estimate C, O, I, H, R, and Q, and find your deficit. Then follow the same idea forward: summarization, retrieval tuning, and sliding-window strategies are all ways of keeping H and R from growing without bound in the first place.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

You count a prompt with one tokenizer and then send it to a different model. Which conclusion best follows from the article?
Question 1 of 2Comparison Reasoning

Focus: Explain why token estimates and context limits should be treated as model- and provider-dependent planning inputs.

A request is 500 tokens over its input allowance. Both old history and low-ranked retrieved passages are candidates to shorten. Which plan best follows the article’s guidance?
Question 2 of 2Scenario Interpretation

Focus: Apply the cut-selection principle by removing low-signal material while preserving useful context and leaving headroom.

References

  1. Effective context engineering for AI agentswww.anthropic.com
  2. Context Window: The Token Budget for Every LLM Callwww.ml4devs.com
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.

A white humanoid toy robot standing on a reflective black surface in a studio setting with a blue and pink gradient background.
beginner
8 min read

A Brief History of LLMs

Large language models look like overnight magic, but they are the visible tip of decades of compounding engineering. If you want to use, trust, and build…

Read tutorial