Skip to content
advanced

LLM KV Cache Memory: Work Through the Inference Math

The weights fit. The model loads. Then the second concurrent request arrives, and the server refuses to admit it.

Published 2026-10-03Updated 2026-10-049 min read
Close-up of rippled sand grains creating a serene desert landscape feel.
Close-up of rippled sand grains creating a serene desert landscape feel. Photo by Ulrick Trappschuh on Pexels.

The weights fit. The model loads. Then the second concurrent request arrives, and the server refuses to admit it.

Nothing about the model changed between the first request and the second. What changed is that the first request is still holding per-token state in device memory, and that state grows with every token it generates. The KV cache is not a speed trick that happens to cost a little memory. It is per-request state that scales with tokens times concurrency, and it competes with the weights for the same pool of GPU memory.

By the end of this article you will be able to estimate that number on paper — before you rent the GPU.

What the Cache Actually Stores

During prefill, the model processes the whole prompt and computes a key vector and a value vector for every token, at every layer. During decode, it computes K and V for exactly one new token per layer per step, appends them, and reuses everything already stored. That reuse is what turns per-step attention cost from recomputing the entire prefix into reading stored state.

The compute saving is real. So is the memory bill.

Three properties matter for the arithmetic:

  • Each layer keeps its own K and V. The cache is not one shared tensor; it is L separate pairs of tensors, one pair per layer.
  • The cache is per-request. Model weights are fixed and shared across every request the server handles. The KV cache is created when a sequence starts, grows token by token, and is discarded when the sequence ends.
  • The cache is not an answer cache. A response or semantic cache stores finished outputs keyed by request similarity. The KV cache stores internal attention state for one in-flight sequence. Same word, different object, different failure modes.

Common mistake: Treating the KV cache as a speed optimization that is "basically free." It is free only while sequences are short and concurrency is low. Past that boundary, it is the constraint.

Knowledge check

Check your understanding

Answer this question before you continue.

Which description matches the KV cache during autoregressive generation?
Misconception Check

Focus: Distinguish per-request KV state from shared model weights and cached answers.

Notation and Assumptions

Before deriving anything, name the symbols:

SymbolMeaning
LLNumber of transformer layers
TTCached tokens for one sequence (prompt plus generated)
BBConcurrent sequences (batch size)
HkvH_{kv}Number of key-value heads
DhD_hHead dimension
bytesBytes per stored element (2 for FP16/BF16, 4 for FP32)

HkvH_{kv} deserves its own line. In multi-head attention, every attention head has its own K and V, so HkvH_{kv} equals the total head count. In grouped-query attention, several query heads share one K/V head, which cuts HkvH_{kv} while leaving the query head count untouched. That single architectural choice is the largest lever on cache size, and it is why the formula cannot always be collapsed into a model-dimension shortcut.

The simplifying assumptions, stated up front:

  • Uniform layer shape — every layer has the same HkvH_{kv} and DhD_h.
  • No sliding-window or sparse attention — every cached token stays resident.
  • No cache quantization — one dtype throughout.
  • No prefix sharing between requests.
  • The cache lives entirely in device memory.
  • TT counts the full cached length, prompt included.

Real serving stacks add allocator overhead, fragmentation, and padding, which push actual allocation above this estimate. Prefix sharing and eviction can push it below. Treat the result as a baseline under the listed assumptions, not a measurement.

Knowledge check

Check your understanding

Answer this question before you continue.

In the article's notation, what does $H_{kv}$ count?
Single Choice

Focus: Interpret the key-value-head count in the simplified cache estimate, including for grouped-query attention.

Deriving the Per-Token Cost

Build the formula one factor at a time.

One token, one layer. The layer stores one key vector and one value vector, each of size Hkv×DhH_{kv} \times D_h elements. That is 2×Hkv×Dh2 \times H_{kv} \times D_h elements per token per layer.

All layers. Each layer keeps its own copy, so multiply by LL:

2×L×Hkv×Dh elements per token2 \times L \times H_{kv} \times D_h \text{ elements per token}

Bytes. Multiply by the dtype width:

bytes per token=2×L×Hkv×Dh×bytes\text{bytes per token} = 2 \times L \times H_{kv} \times D_h \times \text{bytes}

The multi-head simplification. When HkvH_{kv} equals the full head count, the product Hkv×DhH_{kv} \times D_h equals the model dimension dmodeld_{model}, because the heads together span the full representation space. The formula collapses to 2×L×dmodel×bytes2 \times L \times d_{model} \times \text{bytes} per token. This is the form you will see quoted most often.

Where the shortcut breaks. Under grouped-query attention, Hkv<HH_{kv} < H, so Hkv×Dh≠dmodelH_{kv} \times D_h \neq d_{model}. Substituting dmodeld_{model} here overestimates the cache by the grouping ratio. Use the explicit Hkv×DhH_{kv} \times D_h form unless you have confirmed the model is full multi-head.

The full expression. Multiply by sequence length and concurrency:

KV cache memory=2×L×B×T×Hkv×Dh×bytes\text{KV cache memory} = 2 \times L \times B \times T \times H_{kv} \times D_h \times \text{bytes}

Label the factors by what controls them. LL, HkvH_{kv}, and DhD_h are architecture — fixed when you pick the model. TT and BB are workload — set by your traffic. Bytes per element is a deployment choice you make at serving time.

Knowledge check

Check your understanding

Answer this question before you continue.

Under the article's uniform-shape assumptions, which expression gives cache bytes for one token across all layers?
Single Choice

Focus: Derive the per-token cache cost from the stored key and value vectors across all layers.

Worked Example: One Long Request

Take a concrete configuration: L=32L = 32 layers, Hkv=32H_{kv} = 32 heads, Dh=128D_h = 128, and a two-byte dtype.

Bytes per token:

2×32×32×128×2=524,288 bytes≈0.5 MB per token2 \times 32 \times 32 \times 128 \times 2 = 524{,}288 \text{ bytes} \approx 0.5 \text{ MB per token}

At T=4,096T = 4{,}096 tokens, one sequence costs roughly 2 GB2 \text{ GB}. At T=32,768T = 32{,}768, the same sequence costs roughly 16 GB16 \text{ GB} — for a single request.

Now compare that against the weights. A model of this shape has on the order of a few billion parameters; at two bytes per parameter, the weights occupy single-digit gigabytes. The cache for one long-context request is already in the same range. The cache is not a rounding error next to the weights once context gets long — it can rival or exceed them.

Repeat the arithmetic with grouped-query attention, where Hkv=8H_{kv} = 8 instead of 32. Bytes per token drop by a factor of four, to about 0.125 MB0.125 \text{ MB}, and the 32K-token sequence costs roughly 4 GB4 \text{ GB} instead of 16 GB16 \text{ GB}. Same layers, same head dimension, same dtype — a quarter of the memory, purely from the head count.

Note: This is why grouped-query attention shows up in almost every modern serving-oriented model. It is not only a quality-per-compute trade; it is a memory-capacity decision.

Knowledge check

Check your understanding

Answer this question before you continue.

In the worked example, changing $H_{kv}$ from 32 to 8 while keeping the other values and $T=32{,}768$ unchanged makes the one-sequence cache estimate approximately what?
Comparison Reasoning

Focus: Compare the worked example's cache estimate when the number of K/V heads is reduced by grouped-query attention.

Scaling With Concurrency

A two-row comparison shows that a 4K-token sequence uses about 2 GB of cache and allows 30 concurrent sequences, while a 32K-token sequence uses about 16 GB and allows 3.
With a fixed 60 GB cache budget, increasing context length from 4K to 32K tokens cuts the estimated concurrency from 30 sequences to 3.

BB multiplies the entire expression, so concurrency and context length trade directly against each other on a fixed memory budget.

Frame it as a budget equation. Usable device memory, minus weights and runtime overhead, divided by per-sequence cache size, gives the maximum concurrent sequences at a given length:

Bmax=usable memory−weights−overhead2×L×T×Hkv×Dh×bytesB_{max} = \frac{\text{usable memory} - \text{weights} - \text{overhead}}{2 \times L \times T \times H_{kv} \times D_h \times \text{bytes}}

Now carry the earlier numbers through it. Suppose the GPU has 80 GB80 \text{ GB} of usable memory, the weights and runtime overhead consume 20 GB20 \text{ GB}, and each sequence runs at T=4,096T = 4{,}096 tokens. From the worked example, one sequence costs about 2 GB2 \text{ GB}. The remaining budget is 60 GB60 \text{ GB}, so:

Bmax=60 GB2 GB per sequence=30 concurrent sequencesB_{max} = \frac{60 \text{ GB}}{2 \text{ GB per sequence}} = 30 \text{ concurrent sequences}

Thirty requests at 4K tokens each. Push the same budget to T=32,768T = 32{,}768 tokens, where each sequence costs about 16 GB16 \text{ GB}, and the ceiling collapses to three concurrent sequences. Same hardware, same model, same weights — the only variable that moved was context length, and it cut concurrency by a factor of ten.

Run it in reverse and you get the other product decision: fix a target concurrency, solve for the maximum TT that still fits, and that number becomes your effective context limit regardless of what the model card advertises.

This is where batching meets memory. Larger batches raise throughput, but each additional sequence consumes cache, so the batch size is bounded by memory rather than by compute. The failure mode is specific and unpleasant: the scheduler admits one more request, memory runs out mid-generation, and the request either fails outright or forces eviction of another sequence's cache — which means recomputation, which means latency for a user who did nothing wrong.

Where the Simplified Model Breaks

The closed-form estimate is a planning tool, not a measurement. Know where it stops being accurate.

  • Prefix sharing. If many requests carry the same prompt prefix, a serving stack with paged memory layouts can store that prefix once instead of once per request. The naive BB multiplier then overestimates cost for shared-context workloads.
  • Eviction and compression. Dropping or compressing cached entries changes the effective TT or the effective bytes per element. These techniques trade memory against output quality in ways the formula cannot predict.
  • Offloading. Moving cache to host memory or disk relieves capacity pressure but adds transfer latency. It changes the shape of the constraint rather than removing it.
  • Cache quantization. Fewer bytes per element lowers the total, but approximation error compounds across decode steps.
  • Sliding-window and sparse attention. These cap the effective cached length, which is why the linear-in-TT assumption is architecture-dependent, not universal.

Tip: Use the closed-form estimate for capacity planning and admission control. Measure real usage before you trust it as a hard limit.

What to Do With the Number

Read the model configuration to get LL, HkvH_{kv}, and DhD_h rather than guessing. Then decide which constraint you are actually solving for. Fixing context length and deriving concurrency produces a different product decision than fixing concurrency and deriving context length — pick the one that matches how your users behave.

Treat the estimate as an admission-control input. Reject or queue requests that would exceed the budget instead of discovering the limit through out-of-memory failures at 3 a.m.

And recognize when the exercise is overkill. Short-context, low-concurrency workloads rarely need this arithmetic. Optimizing cache memory before you have a concurrency problem is wasted effort.

The cache is per-request state that scales with tokens times concurrency, and that is why long context and high concurrency cannot both be free. Estimate bytes per token from the architecture, multiply by the workload, and treat the result as a budget you enforce — not a curiosity you compute after the outage. Your next concrete step: pull the config.json for the model you actually serve, plug its real layer count and head dimensions into the formula, and compare the result against the memory your serving stack reports. The gap between the two is where your operational margin lives.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

A serving setup has 80 GB of usable device memory; weights and runtime overhead use 20 GB, and each sequence's cache costs about 2 GB. Under the article's simplified budget model, what is the estimated maximum concurrency?
Question 1 of 2Scenario Interpretation

Focus: Use remaining device memory and per-sequence cache cost to estimate a concurrency ceiling.

A workload sends many requests with the same prompt prefix, and its serving stack can store that shared prefix once. What should you expect when comparing actual cache use with the simple estimate that multiplies one sequence's cache by $B$?
Question 2 of 2Scenario Interpretation

Focus: Recognize how prefix sharing can make the simple concurrency-scaled estimate overstate cache use.

References

  1. KV cache offloading | LLM Inference Handbookhandbook.modular.com
  2. KV Cache Memory for LLM Inferencembrenndoerfer.com
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.