Skip to content
intermediate

Calculate the Economics of LLM Cache Hits

A 60% hit rate does not mean a 60% discount. It means you have one number and three missing ones.

Published 2026-10-03Updated 2026-10-049 min read
Close-up of sand grains and small pebbles showing texture and detail.
Close-up of sand grains and small pebbles showing texture and detail. Photo by https://kaboompics.com/ on Pexels.

A 60% hit rate does not mean a 60% discount. It means you have one number and three missing ones.

You have probably seen the scene. A dashboard shows a cache hit rate of 60%, someone in the room says "so we cut costs by 60%," and two weeks later the invoice disagrees. The hit rate was real. The conclusion was not.

Here is the weak model hiding underneath that conclusion: hit rate is a savings number. It is not. Hit rate is one input to an expected-cost equation, and that equation has terms most teams never write down. In this tutorial, we build the equation term by term, run a numerical example, and then stress-test it against the thing that actually kills caching projects: correctness.

Why Hit Rate Is Not a Savings Number

Let me grant the narrow case where the naive belief is roughly true. If you have a pure response cache with zero lookup cost, zero invalidation cost, and no correctness risk — every hit is a byte-for-byte valid answer — then a 60% hit rate really does remove about 60% of your model spend. That world exists mostly in thought experiments.

Real caches leak in two places.

First, a hit still costs something. You paid for the lookup: an embedding call, a vector search, a gateway hop, a storage read. Small, but paid on every request, including misses.

Second, a miss can cost more than baseline. If your cache charges a write premium — many provider-side prompt caches do — then the miss path is not just "the normal request." It is the normal request plus a surcharge for storing the entry.

So the decision variable is not the hit rate you observe on a dashboard. It is expected cost per request, and the gap between the two is exactly the overhead you are paying for the cache's existence.

There is a third leak, and it is the one that matters most: a hit is only a saving if the cached answer is still correct for this request. A cheap wrong answer is not a saving. It is a liability with a discount.

The Four Costs You Need to Name

Before we derive anything, name the terms. Every symbol below maps to a concrete mechanism in a real request path.

C_m — baseline cost of one request with no cache at all. The model call, the tokens, the retrieval, the tool invocations. Whatever your uncached path spends per request.

C_l — cost of the cache lookup itself, per request. The embedding call, the vector search, the gateway hop, the storage read. It is small, but it is not zero, and you pay it on misses too.

h — hit rate, as a fraction of requests served from cache. One warning before we go further: response-cache hits and prompt-cache token shares are different metrics with different economics. A response-cache hit skips the model call entirely. A prompt-cache hit discounts part of one. Do not blend them into a single number, or the equation will lie to you in a way that is hard to detect.

C_i — invalidation and refresh cost, amortized per request. Cache writes, TTL-driven rewrites, re-warming after a flush, and the engineering time to keep entries honest. This is the term teams forget, and it is usually the one that decides whether caching pays.

Picture a single request flowing through the system: it arrives, pays for lookup, then branches. On the hit path it pays the discounted serve cost. On the miss path it pays full price plus its share of the write and refresh work. Every cost above lands at a specific point on that path.

Knowledge check

Check your understanding

Answer this question before you continue.

Which description correctly distinguishes the model's lookup and invalidation terms?
Single Choice

Focus: Identify how lookup cost and invalidation/refresh cost are counted in the per-request model.

Deriving Expected Cost Per Request

A request splits into a cache-hit path labeled h and a miss path labeled 1 − h. The hit serves a cached answer; the miss incurs baseline cost C_m. Both paths lead to shared lookup cost C_l and invalidation cost C_i, then to C_cache = (1 − h)C_m + C_l + C_i.
The hit rate discounts baseline cost on misses; lookup and invalidation costs remain in the total.

Now build the equation from the request path instead of dropping a formula on you.

A request has two outcomes. With probability h, you pay the hit path. With probability (1 − h), you pay the miss path. Expected cost is the probability-weighted sum:

Ccache=h⋅Chit+(1−h)⋅CmissC_{cache} = h \cdot C_{hit} + (1 - h) \cdot C_{miss}

The miss path is the baseline plus the write/refresh share. The hit path is the discounted serve cost plus lookup. Collapse the algebra and you get the working form:

Ccache=(1−h)⋅Cm+Cl+CiC_{cache} = (1 - h) \cdot C_m + C_l + C_i

Read it plainly: you still pay full price on the misses, you always pay for lookup, and you always pay for invalidation. The hit rate only discounts the first term.

Savings are then:

Savings=Cm−CcacheSavings = C_m - C_{cache}

And note the sign. Savings can go negative. If Cl+CiC_l + C_i exceeds h⋅Cmh \cdot C_m, caching costs more than not caching. That is not a hypothetical edge case; it is what happens when you cache a workload that does not repeat.

The assumptions behind this equation are worth stating out loud, because they are where it breaks:

  • Steady-state workload — no cold start dominating the window.
  • Stable entry population — entries are not churning faster than they are reused.
  • No correctness filtering — every hit is treated as valid.
  • Costs that do not change with cache state.

Hold onto that third assumption. We are going to break it deliberately.

Knowledge check

Check your understanding

Answer this question before you continue.

Under the article's expected-cost model, what follows when C_l + C_i is greater than h · C_m?
Misconception Check

Focus: Interpret the expected-cost equation and recognize when caching has negative savings.

Worked Example: $0.010, $0.001, $0.001, h = 0.6

Let us run the numbers end to end. Suppose:

  • C_m = \0.010$ per request
  • C_l = \0.001$ per request
  • C_i = \0.001$ per request
  • h=0.6h = 0.6

Miss path: (1 - 0.6) \times \0.010 = $0.004$.

Add lookup and invalidation: \0.004 + $0.001 + $0.001 = $0.006$.

So C_{cache} = \0.006perrequest,andsavingsareper request, and savings are$0.010 - $0.006 = $0.004$ per request.

That is 40%, not 60%. The headline hit rate overstated the saving by a third.

Now push on the hit rate, because this is where the economics get uncomfortable.

Hit rate hC_cacheSavings per request
0.6$0.006$0.004
0.3$0.009$0.001
0.1$0.011−$0.001

At h = 0.3, savings collapse to a tenth of a cent. At h = 0.1, caching is underwater. The break-even hit rate is:

h∗=Cl+CiCm=0.001+0.0010.010=0.2h^* = \frac{C_l + C_i}{C_m} = \frac{0.001 + 0.001}{0.010} = 0.2

Below a 20% hit rate, this cache destroys value. That is the number I would put on the wall before anyone promises a finance team a percentage.

Knowledge check

Check your understanding

Answer this question before you continue.

Using the example's C_m = $0.010, C_l = $0.001, C_i = $0.001, and h = 0.6, what are the expected cached cost and savings per request?
Output Prediction

Focus: Calculate cached cost and savings for the article's worked numerical example.

Where the Equation Breaks: Freshness and Personalization

Everything above assumes every hit is a valid hit. That assumption is the one most likely to be false in production.

A hit is only a saving if the cached answer is still correct for this request. Stale entries and improperly personalized entries convert a cost saving into a correctness incident. The arithmetic does not care; your users will.

You have two honest ways to account for this:

Bypass unsafe hits. If the entry is stale or belongs to a different user's context, treat the request as a miss. Effective hit rate drops, and the equation above still holds — with a smaller h.

Count unsafe hits as risk. Multiply the probability of serving a bad answer by the expected cost of that answer. In most applications, one confidently wrong answer delivered at scale dwarfs the token saving on the request that produced it.

Concrete failure shapes worth recognizing:

  • A price or policy that changed since the entry was written.
  • A cached answer served into a different user's context.
  • A semantic cache that matched on surface similarity but not intent.

Note the asymmetry. A 40% saving on tokens is small next to one wrong answer delivered confidently to a thousand users. The cache does not know the difference. You have to.

Adjusting the Model for Safe Hits Only

Introduce h_safe — the fraction of requests where a hit is both available and safe to serve. Replace h with h_safe in the cost equation, and the worked example changes character.

With hsafe=0.4h_{safe} = 0.4 instead of 0.6:

Ccache=0.6×$0.010+$0.001+$0.001=$0.008C_{cache} = 0.6 \times \$0.010 + \$0.001 + \$0.001 = \$0.008

Savings: \0.002$ per request. The safety filter halved the benefit.

And here is the part that makes this a real tradeoff rather than a one-time adjustment: raising h_safe usually costs money. Shorter TTLs, narrower cache keys, per-user partitioning, more aggressive invalidation — every one of those pushes CiC_i up. The variables move against each other. Hit rate, freshness, personalization, and invalidation cost form a surface, and the equation is how you see which direction is winning.

Decision rule: If you cannot state h_safe and C_i from measured data, you do not have a caching business case. You have a hypothesis.

Knowledge check

Check your understanding

Answer this question before you continue.

For the article's example with h_safe = 0.4, C_m = $0.010, C_l = $0.001, and C_i = $0.001, what does the safe-hit model predict?
Scenario Interpretation

Focus: Recompute expected cost and savings when only safe hits count toward the hit rate.

What to Measure Before You Commit

The theory is only useful if you replace its inputs with observed ones. Here is the short measurement plan.

Measure C_m from real uncached requests, not list prices. Batch discounts, output-heavy workloads, and retrieval costs all move it. Your baseline is what you actually spend, not what the pricing page says.

Measure h and h_safe separately. A single blended hit rate hides the correctness filter and flatters the result. If you only track one number, you will systematically overestimate savings.

Track C_l and C_i as first-class line items. Include the engineering time spent on invalidation. If you leave it out, the equation quietly lies in your favor.

Watch for regressions. A prompt template change that breaks prefix stability can halve the hit rate and show up later as a cost spike. Alert on per-workload drops, not fleet averages.

Common mistake: Comparing a cached path against a baseline you never actually measured. If your "before" number came from a pricing page and your "after" number came from a bill, you are not comparing two systems. You are comparing two stories.

The Next Step

Write the four costs down. Compute your break-even hit rate with h∗=(Cl+Ci)/Cmh^* = (C_l + C_i) / C_m. Then subtract the hits you would not dare serve.

Pick one workload this week. Instrument it, measure CmC_m and hsafeh_{safe} for seven days, and re-run the arithmetic with real numbers. If the saving survives that, you have a business case. If it does not, you just saved yourself from shipping a cache that was quietly charging rent.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

A cache reports many hits, but some entries may be stale or belong to another user's context. Which treatment matches the article's guidance?
Question 1 of 2Scenario Interpretation

Focus: Choose how to account for stale or improperly personalized cache hits in an economic estimate.

A team uses a pricing-page estimate for C_m, reports only blended hit rate h, and omits invalidation engineering time. What is the strongest improvement to its evaluation?
Question 2 of 2Comparison Reasoning

Focus: Select measurements needed to evaluate whether caching produces a credible business case.

References

  1. Cache Hit Rate (LLM): Definition, SQL & How to Track It in Metabasewww.metabase.com
  2. Pricing - Claude Platform Docsdocs.anthropic.com
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.