Calculate the Economics of LLM Cache Hits
A 60% hit rate does not mean a 60% discount. It means you have one number and three missing ones.

Key topics
A 60% hit rate does not mean a 60% discount. It means you have one number and three missing ones.
You have probably seen the scene. A dashboard shows a cache hit rate of 60%, someone in the room says "so we cut costs by 60%," and two weeks later the invoice disagrees. The hit rate was real. The conclusion was not.
Here is the weak model hiding underneath that conclusion: hit rate is a savings number. It is not. Hit rate is one input to an expected-cost equation, and that equation has terms most teams never write down. In this tutorial, we build the equation term by term, run a numerical example, and then stress-test it against the thing that actually kills caching projects: correctness.
Why Hit Rate Is Not a Savings Number
Let me grant the narrow case where the naive belief is roughly true. If you have a pure response cache with zero lookup cost, zero invalidation cost, and no correctness risk — every hit is a byte-for-byte valid answer — then a 60% hit rate really does remove about 60% of your model spend. That world exists mostly in thought experiments.
Real caches leak in two places.
First, a hit still costs something. You paid for the lookup: an embedding call, a vector search, a gateway hop, a storage read. Small, but paid on every request, including misses.
Second, a miss can cost more than baseline. If your cache charges a write premium — many provider-side prompt caches do — then the miss path is not just "the normal request." It is the normal request plus a surcharge for storing the entry.
So the decision variable is not the hit rate you observe on a dashboard. It is expected cost per request, and the gap between the two is exactly the overhead you are paying for the cache's existence.
There is a third leak, and it is the one that matters most: a hit is only a saving if the cached answer is still correct for this request. A cheap wrong answer is not a saving. It is a liability with a discount.
The Four Costs You Need to Name
Before we derive anything, name the terms. Every symbol below maps to a concrete mechanism in a real request path.
C_m — baseline cost of one request with no cache at all. The model call, the tokens, the retrieval, the tool invocations. Whatever your uncached path spends per request.
C_l — cost of the cache lookup itself, per request. The embedding call, the vector search, the gateway hop, the storage read. It is small, but it is not zero, and you pay it on misses too.
h — hit rate, as a fraction of requests served from cache. One warning before we go further: response-cache hits and prompt-cache token shares are different metrics with different economics. A response-cache hit skips the model call entirely. A prompt-cache hit discounts part of one. Do not blend them into a single number, or the equation will lie to you in a way that is hard to detect.
C_i — invalidation and refresh cost, amortized per request. Cache writes, TTL-driven rewrites, re-warming after a flush, and the engineering time to keep entries honest. This is the term teams forget, and it is usually the one that decides whether caching pays.
Picture a single request flowing through the system: it arrives, pays for lookup, then branches. On the hit path it pays the discounted serve cost. On the miss path it pays full price plus its share of the write and refresh work. Every cost above lands at a specific point on that path.
Knowledge check
Check your understanding
Answer this question before you continue.
Deriving Expected Cost Per Request
Now build the equation from the request path instead of dropping a formula on you.
A request has two outcomes. With probability h, you pay the hit path. With probability (1 − h), you pay the miss path. Expected cost is the probability-weighted sum:
The miss path is the baseline plus the write/refresh share. The hit path is the discounted serve cost plus lookup. Collapse the algebra and you get the working form:
Read it plainly: you still pay full price on the misses, you always pay for lookup, and you always pay for invalidation. The hit rate only discounts the first term.
Savings are then:
And note the sign. Savings can go negative. If exceeds , caching costs more than not caching. That is not a hypothetical edge case; it is what happens when you cache a workload that does not repeat.
The assumptions behind this equation are worth stating out loud, because they are where it breaks:
- Steady-state workload — no cold start dominating the window.
- Stable entry population — entries are not churning faster than they are reused.
- No correctness filtering — every hit is treated as valid.
- Costs that do not change with cache state.
Hold onto that third assumption. We are going to break it deliberately.
Knowledge check
Check your understanding
Answer this question before you continue.
Worked Example: $0.010, $0.001, $0.001, h = 0.6
Let us run the numbers end to end. Suppose:
- C_m = \0.010$ per request
- C_l = \0.001$ per request
- C_i = \0.001$ per request
Miss path: (1 - 0.6) \times \0.010 = $0.004$.
Add lookup and invalidation: \0.004 + $0.001 + $0.001 = $0.006$.
So C_{cache} = \0.006$0.010 - $0.006 = $0.004$ per request.
That is 40%, not 60%. The headline hit rate overstated the saving by a third.
Now push on the hit rate, because this is where the economics get uncomfortable.
| Hit rate h | C_cache | Savings per request |
|---|---|---|
| 0.6 | $0.006 | $0.004 |
| 0.3 | $0.009 | $0.001 |
| 0.1 | $0.011 | −$0.001 |
At h = 0.3, savings collapse to a tenth of a cent. At h = 0.1, caching is underwater. The break-even hit rate is:
Below a 20% hit rate, this cache destroys value. That is the number I would put on the wall before anyone promises a finance team a percentage.
Knowledge check
Check your understanding
Answer this question before you continue.
Where the Equation Breaks: Freshness and Personalization
Everything above assumes every hit is a valid hit. That assumption is the one most likely to be false in production.
A hit is only a saving if the cached answer is still correct for this request. Stale entries and improperly personalized entries convert a cost saving into a correctness incident. The arithmetic does not care; your users will.
You have two honest ways to account for this:
Bypass unsafe hits. If the entry is stale or belongs to a different user's context, treat the request as a miss. Effective hit rate drops, and the equation above still holds — with a smaller h.
Count unsafe hits as risk. Multiply the probability of serving a bad answer by the expected cost of that answer. In most applications, one confidently wrong answer delivered at scale dwarfs the token saving on the request that produced it.
Concrete failure shapes worth recognizing:
- A price or policy that changed since the entry was written.
- A cached answer served into a different user's context.
- A semantic cache that matched on surface similarity but not intent.
Note the asymmetry. A 40% saving on tokens is small next to one wrong answer delivered confidently to a thousand users. The cache does not know the difference. You have to.
Adjusting the Model for Safe Hits Only
Introduce h_safe — the fraction of requests where a hit is both available and safe to serve. Replace h with h_safe in the cost equation, and the worked example changes character.
With instead of 0.6:
Savings: \0.002$ per request. The safety filter halved the benefit.
And here is the part that makes this a real tradeoff rather than a one-time adjustment: raising h_safe usually costs money. Shorter TTLs, narrower cache keys, per-user partitioning, more aggressive invalidation — every one of those pushes up. The variables move against each other. Hit rate, freshness, personalization, and invalidation cost form a surface, and the equation is how you see which direction is winning.
Decision rule: If you cannot state h_safe and C_i from measured data, you do not have a caching business case. You have a hypothesis.
Knowledge check
Check your understanding
Answer this question before you continue.
What to Measure Before You Commit
The theory is only useful if you replace its inputs with observed ones. Here is the short measurement plan.
Measure C_m from real uncached requests, not list prices. Batch discounts, output-heavy workloads, and retrieval costs all move it. Your baseline is what you actually spend, not what the pricing page says.
Measure h and h_safe separately. A single blended hit rate hides the correctness filter and flatters the result. If you only track one number, you will systematically overestimate savings.
Track C_l and C_i as first-class line items. Include the engineering time spent on invalidation. If you leave it out, the equation quietly lies in your favor.
Watch for regressions. A prompt template change that breaks prefix stability can halve the hit rate and show up later as a cost spike. Alert on per-workload drops, not fleet averages.
Common mistake: Comparing a cached path against a baseline you never actually measured. If your "before" number came from a pricing page and your "after" number came from a bill, you are not comparing two systems. You are comparing two stories.
The Next Step
Write the four costs down. Compute your break-even hit rate with . Then subtract the hits you would not dare serve.
Pick one workload this week. Instrument it, measure and for seven days, and re-run the arithmetic with real numbers. If the saving survives that, you have a business case. If it does not, you just saved yourself from shipping a cache that was quietly charging rent.
Knowledge check
Final check
Finish the article by checking the ideas you just learned.
References
Want a more structured LLMOps path?
Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.
Large Language Models Starter Pack
A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.
- 227-page Illustrated PDF edition
- 12 guided LLM engineering chapters
- Visual concept diagrams
- Self-assessment quizzes
- Bonus deep-dive sections
- Prompt design, structured output, context windows & RAG pipelines
- Agents, tool calling, prompt injection, evaluation & application lifecycles
Coming soon


