Skip to content
intermediate

How to Aggregate LLM Latency and Usage: A Worked Example

Your dashboard says the average response takes 1.2 seconds. Your users say the app feels slow. Both statements are true, and that is the problem.

Published 2026-10-03Updated 2026-10-0413 min read
Expansive desert landscape featuring majestic sand dunes under a clear blue sky, perfect for travel and nature themes.
Expansive desert landscape featuring majestic sand dunes under a clear blue sky, perfect for travel and nature themes. Photo by Guy Seela on Pexels.

Your dashboard says the average response takes 1.2 seconds. Your users say the app feels slow. Both statements are true, and that is the problem.

A metric is not a property of your system. It is a property of a decision you made: which events you counted, over what window, and which summary you applied. Change any one of those three and the number changes — honestly, defensibly, and in a way that makes two teams argue about the same week of traffic.

This article is about that decision layer. We will take ten hypothetical requests, compute the summaries by hand, and watch how the same trace produces two different "average latencies" depending on what we decided to count. Aggregation is a definition problem before it is a math problem.

Why Your Latency Number Is a Definition, Not a Measurement

Every aggregate you report is the output of a function:

metric = summarize(events, window, statistic)

Three inputs, one number. If two dashboards disagree, they almost always differ on one of those three inputs, not on the arithmetic.

The failure families split cleanly:

  • Scope errors — you counted the wrong events. Retries logged as new requests. Timeouts that never emitted a completion event. Cached responses mixed with fresh generations.
  • Summary errors — you counted the right events but described them with the wrong statistic. A mean standing in for a distribution with a long right tail.

Traces give you raw events. Aggregation is the layer that turns those events into claims about your system. If you have not yet set up tracing, that is the prerequisite — this article assumes you already have request events landing somewhere queryable and you can read a span.

The rest of this article is a single worked example that makes both failure families visible.

Notation and Scope: What Counts as One Event

Before any arithmetic, fix the notation. For a single request rr:

SymbolMeaning
tstartt_{\text{start}}When the request was sent
tfirstt_{\text{first}}When the first output token arrived
tendt_{\text{end}}When the last output token arrived
ninn_{\text{in}}Input tokens
noutn_{\text{out}}Output tokens

From those, the derived per-request quantities:

  • TTFT (time to first token): tfirst−tstartt_{\text{first}} - t_{\text{start}}
  • Generation duration: tend−tfirstt_{\text{end}} - t_{\text{first}}
  • End-to-end latency: tend−tstartt_{\text{end}} - t_{\text{start}}
  • Tokens per second: nout/generation durationn_{\text{out}} / \text{generation duration}

End-to-end latency decomposes as TTFT+generation duration\text{TTFT} + \text{generation duration}. That decomposition matters later: a slow response is either slow to start, slow to stream, or simply long.

Now the scope rules — the decisions that determine which events enter the set:

  1. What is one request? A retry after a timeout is a second model call but arguably one logical user request. A fallback to a different provider is the same story.
  2. Do cached responses count? A cache hit has near-zero latency and zero output tokens. Including it drags every summary toward optimism.
  3. Do failed requests count? A request that timed out at 30 seconds and never completed is a real user experience. Excluding it biases the tail downward.
  4. Are units and windows comparable? Milliseconds versus seconds, and a five-minute window versus a week, produce different numbers from identical events.

Common mistake: Logging each retry as a new request. The denominator inflates, the retry's latency enters the distribution as if it were a fresh user request, and the tail looks healthier than the experience actually was. If a request retried twice before succeeding, you have one slow user experience and three telemetry events.

The assumption that makes any aggregation valid: comparable events, comparable units, comparable windows. Break it and the number is not wrong in the arithmetic — it is wrong in the claim.

Knowledge check

Check your understanding

Answer this question before you continue.

A logical user request times out, retries twice, and then succeeds. Under the article’s distinction, which description is accurate?
Scenario Interpretation

Focus: Distinguish logical user requests from model-call telemetry events when retries occur.

A Worked Trace: Ten Requests, Two Metrics

Two horizontal bars compare requests 1 and 10. Each bar is split into TTFT and generation duration: request 1 is 0.30 plus 1.20 seconds, totaling 1.50 seconds; request 10 is 1.80 plus 2.50 seconds, totaling 4.30 seconds. Both have a 100 tokens-per-second generation rate, while request 10 produces more output tokens.
The decomposition shows that request 10 takes longer because it starts later and generates more tokens—not because its streaming rate is lower.

Here is a compact trace. Ten requests, all to the same model version, all with the same prompt template, all in one hour. TTFT and end-to-end latency in seconds.

#TTFTGen durationEnd-to-endninn_{\text{in}}noutn_{\text{out}}
10.301.201.50400120
20.280.901.1838090
30.351.601.95420160
40.311.101.41390110
50.290.801.0941080
60.331.401.73400140
70.301.001.30395100
80.270.700.9738570
90.341.301.64405130
101.802.504.30400250

Request 10 is the outlier: a cold start, a provider hiccup, or a genuinely longer answer. We will keep it, because excluding it is exactly the kind of scope decision that hides user pain.

Per-request tokens per second. For request 1: 120/1.20=100120 / 1.20 = 100 tokens/s. For request 10: 250/2.50=100250 / 2.50 = 100 tokens/s. Same streaming rate — request 10 is slow because it started slowly and produced more tokens, not because generation degraded.

Arithmetic mean of end-to-end latency:

Lˉ=1.50+1.18+1.95+1.41+1.09+1.73+1.30+0.97+1.64+4.3010=17.0710=1.707 s\bar{L} = \frac{1.50 + 1.18 + 1.95 + 1.41 + 1.09 + 1.73 + 1.30 + 0.97 + 1.64 + 4.30}{10} = \frac{17.07}{10} = 1.707\text{ s}

Arithmetic mean of TTFT:

TTFT‾=0.30+0.28+0.35+0.31+0.29+0.33+0.30+0.27+0.34+1.8010=4.5710=0.457 s\overline{\text{TTFT}} = \frac{0.30 + 0.28 + 0.35 + 0.31 + 0.29 + 0.33 + 0.30 + 0.27 + 0.34 + 1.80}{10} = \frac{4.57}{10} = 0.457\text{ s}

Notice what request 10 did: it pulled the mean TTFT from roughly 0.31 s to 0.457 s. One request out of ten moved the average by nearly 50%.

Percentiles. Sort the end-to-end latencies ascending:

0.97, 1.09, 1.18, 1.30, 1.41, 1.50, 1.64, 1.73, 1.95, 4.30

Using the nearest-rank convention — the smallest value at or above the kk-th percentile position — with N=10N = 10:

  • p50 (median): position ⌈0.50×10⌉=5\lceil 0.50 \times 10 \rceil = 5 → 1.41 s
  • p90: position ⌈0.90×10⌉=9\lceil 0.90 \times 10 \rceil = 9 → 1.95 s
  • p95: position ⌈0.95×10⌉=10\lceil 0.95 \times 10 \rceil = 10 → 4.30 s

That p95 is worth pausing on. With ten samples, the 95th percentile is the worst request. Small samples make tail percentiles jumpy — a fact that will matter when you compare two weeks of traffic.

Note: Percentile conventions differ. Nearest-rank picks an actual observed value. Linear interpolation (the "inclusive" method common in monitoring tools) averages between neighbors. On this trace, interpolated p95 lands near 3.13 s instead of 4.30 s. Same events, same percentile label, different number. Always state the convention.

Token usage. Total input tokens: 3,985. Total output tokens: 1,250. Mean input per request: 398.5. Mean output per request: 125.

Now the weighting question, and it is worth being precise because the labels get sloppy. The mean output tokens per request is the total divided by the request count: 1250/10=1251250 / 10 = 125. That is a request-weighted average — every request contributes one data point, regardless of how long it was.

A token-weighted rate answers a different question: across all the tokens you generated, how fast did they stream? You compute it as aggregate output tokens divided by aggregate generation time:

token-weighted rate=∑nout∑generation duration=125012.50=100 tokens/s\text{token-weighted rate} = \frac{\sum n_{\text{out}}}{\sum \text{generation duration}} = \frac{1250}{12.50} = 100 \text{ tokens/s}

The request-weighted mean of per-request rates is also 100 tokens/s here, because every request streamed at the same rate. Change one request and the two diverge. Suppose request 10 had streamed at 50 tokens/s instead of 100 — its generation duration would be 250/50=5.0250 / 50 = 5.0 s. The request-weighted mean of per-request rates becomes (9×100+50)/10=95(9 \times 100 + 50) / 10 = 95 tokens/s. The token-weighted rate becomes 1250/(9×1.0+5.0)=1250/14.0≈89.31250 / (9 \times 1.0 + 5.0) = 1250 / 14.0 \approx 89.3 tokens/s. The token-weighted number is lower because the slow request carried more tokens, and those tokens dominate the aggregate.

Neither number is wrong. They answer different questions. Request-weighted tells you what a typical request experienced. Token-weighted tells you how efficiently the serving layer moved the total workload.

The same trace, two summaries. If you report "average latency: 1.71 s," you are describing a system where the typical request finishes in about 1.4 s and one request in ten takes over four seconds. If you report "p50: 1.41 s, p95: 4.30 s," you are describing the same system honestly. The first number is not wrong. It is just answering a different question than the one your users are asking.

Knowledge check

Check your understanding

Answer this question before you continue.

Using the article’s nearest-rank convention, what is the p90 end-to-end latency for the ten-request trace?
Output Prediction

Focus: Apply nearest-rank percentile positioning to the worked trace’s sorted end-to-end latencies.

Average vs Tail Percentiles: What Each One Hides

Latency distributions are right-skewed. Most requests cluster near the median; a few stretch far to the right. The mean is sensitive to those few. The median barely notices them.

That asymmetry is why "healthy average, unhappy users" is such a common pairing. On our trace, the mean sits at 1.71 s — comfortably under a two-second budget. The p95 sits at 4.30 s, well over it. If your timeout is three seconds, roughly one request in ten is failing, and the average never told you.

The tail also compounds. A request that calls the model three times in a pipeline — retrieve, generate, verify — inherits three independent chances to hit the slow path. If each call has a 5% chance of being slow, the pipeline has roughly a 1−0.953≈14%1 - 0.95^3 \approx 14\% chance of containing at least one slow call. Tail latency in multi-step systems is not the tail of one call. It is the tail of the slowest call in the chain.

My rule for choosing a summary:

  • p50 for typical user experience. This is what most people feel most of the time.
  • p95 and p99 for user-facing promises, timeouts, and SLOs. This is where perceived slowness lives.
  • The mean only when you also report the spread. A mean without a p95 is a claim without a caveat.

Tip: If you can only put one latency number on a dashboard, put p95 there. It is the number that predicts complaints.

Knowledge check

Check your understanding

Answer this question before you continue.

In the worked trace, the mean latency is below two seconds while p95 is 4.30 seconds. What conclusion is best supported?
Misconception Check

Focus: Explain why a favorable mean latency does not establish that tail latency meets a user-facing budget.

Aggregating Usage Without Double-Counting

Token aggregation looks like addition. It is actually another scope problem.

Input, output, and cached tokens are different quantities with different costs. Collapsing them into a single "tokens" number destroys the signal you need for cost analysis — cached input tokens are typically far cheaper than fresh output tokens, and a system that shifts work from output to cached input can look flat on total tokens while its bill drops.

Then there is the retry problem again. If a request retries twice, you have three sets of token counts. Are you counting billed tokens (what the provider charged you) or logical-request tokens (what one user interaction consumed)? Both are legitimate. They answer different questions. Mixing them produces a number that answers neither.

The same weighting question from latency applies here:

QuestionWeightingWhy
What does a typical request cost?Request-weightedEach user interaction counts once
How efficient is the serving layer?Token-weightedLong requests dominate the work
What will the bill be?Sum of billed tokensProvider charges per token, not per request

A long request and a short request should not count equally when you care about throughput. They should count equally when you care about user experience.

Practical rule: Store raw per-request usage events and derive aggregates from them. If you only store the aggregate, you can never re-slice by model version, prompt version, or input-token bucket — and you will eventually need to.

How Missing and Differently Scoped Events Distort Comparisons

This is where the definition problem becomes expensive. Two weeks of data, two dashboards, two conclusions.

Missing events. Dropped traces, sampled telemetry, and timeouts that never emit a completion event all bias the tail downward. A request that hung for 30 seconds and was killed by a client timeout may never produce a completion span. Your p95 improves because the worst requests vanished from the dataset, not because the system got faster.

Scope drift. A deploy adds a retry, a cache, or a streaming path. The events still flow. The denominator silently changes meaning. Week-over-week latency "improves" because cached responses now dilute the distribution.

Segment before comparing. Model version, prompt version, input-token bucket, endpoint. A TTFT regression is often a prompt that grew, not a model that slowed — check input size before blaming infrastructure. On our trace, request 10's TTFT of 1.80 s would look alarming until you noticed its input was the same size as the others, which rules out prompt growth and points at a cold start or provider event.

The detection habit. Compare event counts and token totals alongside latency. A latency drop accompanied by a falling event count is usually a measurement change, not a performance win. If your p95 improved 20% and your request count dropped 20%, you did not get faster. You lost data.

Warning: Never compare percentiles computed with different conventions or over different event scopes. Nearest-rank p95 from one tool and interpolated p95 from another are not the same statistic, even when they share a label.

Knowledge check

Check your understanding

Answer this question before you continue.

A dashboard’s p95 improves by 20% from one week to the next, but its request count also drops by 20%. Which interpretation best follows the article?
Scenario Interpretation

Focus: Recognize that falling event counts can make latency appear to improve when data is missing.

When to Use Which Summary

A short decision table, because the choice should be mechanical once you know the question:

You want to know...UseAvoid
Typical user experiencep50Mean
Whether you are meeting a promisep95 / p99Mean
Total capacity and costMean and sumsPercentiles alone
Serving efficiencyToken-weighted throughputRequest-weighted
Whether a change helpedSegmented p50 + p95, same scopeBlended numbers

Two things not to do:

  • Do not blend across models, endpoints, or prompt versions when you intend to compare them. A single number across three model versions tells you nothing about any of them.
  • Do not compare percentiles across different conventions or scopes. If the scope changed, the comparison is invalid before the arithmetic starts.

Write Down the Definition Before You Trust the Number

Before you trust any latency or usage number, write down three things: the event scope, the time window, and the percentile convention that produced it. If you cannot fill in all three, you do not have a metric. You have a number that happens to be on a screen.

Here is the next move I would make. Pick one metric on an existing dashboard — the average latency, the p95, the daily token total — and recompute it by hand from the raw events for a single hour. Then compare. If your hand-computed number differs from the dashboard, you have found a scope decision nobody documented. That gap is the most valuable thing you will learn this week, and it costs one query and ten minutes.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

Suppose nine requests each generate 100 tokens at 100 tokens/s, while a tenth generates 250 tokens at 50 tokens/s. Which pair correctly describes the request-weighted mean of per-request rates and the token-weighted aggregate rate?
Question 1 of 2Comparison Reasoning

Focus: Distinguish request-weighted mean generation rate from token-weighted aggregate throughput.

Two teams want to compare their reported p95 latency. Which set of details should they align or document to make the comparison meaningful?
Question 2 of 2Single Choice

Focus: Identify the definitions needed to make latency summaries meaningfully comparable.

References

  1. Understand LLM latency and throughput metrics | Anyscale Docsdocs.anyscale.com
  2. Trace an LLM application tutorial - Docs by LangChaindocs.langchain.com
  3. Metrics - Pipecatdocs.pipecat.ai
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.