Skip to content
intermediate

Model LLM Cost, Latency, and Quality with Weighted Utility

Three engineers, three candidates, three different winners. The meeting ends the way it always does: whoever speaks last wins, and the decision gets…

Published 2026-10-03Updated 2026-10-0410 min read
Close-up of rippled sand dunes creating a natural wavy texture in a barren desert landscape.
Close-up of rippled sand dunes creating a natural wavy texture in a barren desert landscape. Photo by Frederic Hancke on Pexels.

Three engineers, three candidates, three different winners. The meeting ends the way it always does: whoever speaks last wins, and the decision gets relitigated next sprint.

The tradeoff itself is not the problem. You already know that cheaper models tend to be faster and weaker, that bigger models tend to be slower and stronger, and that your token volume and request patterns decide how much any of that matters. That intuition is settled ground. What is missing is a way to combine three incomparable numbers into one comparable number — and to show your work while doing it.

A weighted utility score does not remove judgment from the decision. It forces judgment into the open, where it can be argued with, tested, and revised.

Why "Best Model" Is an Unfinished Sentence

When someone says "model A is best," they have silently finished a sentence that started with "given my priorities." The failure in the meeting is not that people disagree about facts. It is that each stakeholder is optimizing a different objective and calling it quality.

So start by naming the three raw metrics and their directions:

  • Quality qq — higher is better. It must come from a defined evaluation, not vibes.
  • Latency ll — lower is better.
  • Cost cc — lower is better.

Direction matters because it determines the sign of every normalization you are about to write. Get a sign wrong and your "best" candidate becomes your worst.

The deliverable is a scalar UU per candidate, plus the ability to re-run it when assumptions move. That second part is the real product. A ranking you cannot re-derive is an opinion wearing a table.

Notation, Assumptions, and What the Score Can't Do

Before any arithmetic, define the pieces.

A candidate is one configuration — a model plus its settings — carrying a triple (q,l,c)(q, l, c) measured on a consistent basis. If one candidate's quality comes from a benchmark and another's comes from a colleague's impression, you are not comparing candidates. You are comparing measurement methods.

The bounds are qmin⁡q_{\min}, qmax⁡q_{\max}, lmin⁡l_{\min}, lmax⁡l_{\max}, cmin⁡c_{\min}, cmax⁡c_{\max}. These are chosen, not discovered. But there are two distinct ways to choose them, and mixing them up will quietly corrupt your comparison.

Candidate-derived bounds use the worst and best values in your current candidate set. If your three candidates have quality 8, 7, and 6, then qmin⁡=6q_{\min}=6 and qmax⁡=8q_{\max}=8. This is the easiest policy, and it makes the best candidate score exactly 1.0 on that metric. The catch: add a fourth candidate at quality 10, and every existing quality utility shrinks. The frame moved, so the scores moved.

Fixed decision-range bounds come from outside the candidate set — a quality floor your product requires, a latency ceiling your users tolerate, a cost ceiling your budget allows. If you decide that quality 10 is the best value worth caring about and quality 6 is the worst acceptable, those bounds stay put even when the candidate set changes. A candidate at quality 8 then scores 0.5, not 1.0, because it occupies only part of the range you defined.

Both policies are legitimate. The mistake is switching between them mid-analysis, or reporting a score without saying which one produced it. For the rest of this article, the worked example uses candidate-derived bounds, and the bounds-sensitivity section switches to a fixed anchor on purpose — so you can see the difference.

The model rests on four assumptions:

  1. Metrics are commensurable after normalization.
  2. Preferences are stable across the decision.
  3. Weights are nonnegative and sum to one.
  4. The objectives are substitutable rather than hard constraints.

That fourth assumption is the one that bites. Linear additive utility assumes you will trade a little quality for a little latency at a constant rate. Often you won't — and we will break the model on purpose later to show where.

Note: Under candidate-derived bounds, the score is relative to the candidate set, not absolute. Add a new candidate and every existing score can shift. Under fixed bounds, scores stay anchored but may never reach 1.0.

Knowledge check

Check your understanding

Answer this question before you continue.

With candidate-derived quality bounds, the current candidates have quality values 6, 7, and 8. What happens to the quality utility of the candidate at 8 if a candidate at 10 is added, while the minimum remains 6?
Scenario Interpretation

Focus: Explain how candidate-derived normalization bounds respond when the comparison set changes.

Deriving the Normalized Utilities

Quality, latency, and cost feed through normalization into utilities from zero to one. Quality increases toward its best value, while latency and cost are inverted so lower values score higher. The three utilities are multiplied by their respective weights and combined into one utility score U.
Normalization puts unlike metrics on a shared scale; the weights then make stakeholder priorities explicit in the combined score.

Raw units are incomparable. Milliseconds, dollars, and a 1–5 quality rating cannot be added. Normalization maps each metric onto a common [0,1][0, 1] scale where 0 is the worst value in the comparison range and 1 is the best.

Quality rises with the raw value, so the numerator measures distance above the floor:

uq=q−qmin⁡qmax⁡−qmin⁡u_q = \frac{q - q_{\min}}{q_{\max} - q_{\min}}

The denominator is the width of the comparison range. A candidate at the top of the range gets 1; at the bottom, 0.

Latency falls as the raw value rises, so the numerator is inverted. Fast must map to 1, not 0:

ul=lmax⁡−llmax⁡−lmin⁡u_l = \frac{l_{\max} - l}{l_{\max} - l_{\min}}

Cost uses the same inversion for the same reason:

uc=cmax⁡−ccmax⁡−cmin⁡u_c = \frac{c_{\max} - c}{c_{\max} - c_{\min}}

Every utility lands in [0,1][0, 1]. Now combine them:

U=wquq+wlul+wcuc,w≥0,wq+wl+wc=1U = w_q u_q + w_l u_l + w_c u_c, \quad w \geq 0, \quad w_q + w_l + w_c = 1

The sum-to-one constraint is what makes the weights readable. wl=0.5w_l = 0.5 means latency carries half of the total preference. Without the constraint, weights are arbitrary scale factors and you cannot compare two weight vectors at all.

Knowledge check

Check your understanding

Answer this question before you continue.

Latency is lower-is-better. Using bounds of 150 ms and 900 ms, what normalized latency utility does a candidate with latency 300 ms receive?
Single Choice

Focus: Apply the inverted normalization for a lower-is-better metric.

Worked Example: Three Candidates, One Ranking

Three hypothetical candidates, small numbers you can check by hand:

CandidateQuality qqLatency ll (ms)Cost cc (per 1K requests)
A8.0900$4.00
B7.0300$2.00
C6.0150$0.50

Candidate-derived bounds: q∈[6,8]q \in [6, 8], l∈[150,900]l \in [150, 900], c∈[0.50,4.00]c \in [0.50, 4.00].

Normalize each metric:

Candidateuqu_qulu_lucu_c
A(8−6)/(8−6)=1.00(8-6)/(8-6) = 1.00(900−900)/(900−150)=0.00(900-900)/(900-150) = 0.00(4−4)/(4−0.5)=0.00(4-4)/(4-0.5) = 0.00
B(7−6)/2=0.50(7-6)/2 = 0.50(900−300)/750=0.80(900-300)/750 = 0.80(4−2)/3.5=0.57(4-2)/3.5 = 0.57
C(6−6)/2=0.00(6-6)/2 = 0.00(900−150)/750=1.00(900-150)/750 = 1.00(4−0.5)/3.5=1.00(4-0.5)/3.5 = 1.00

Now pick a starting weight vector for a latency-sensitive product: wq=0.3w_q = 0.3, wl=0.5w_l = 0.5, wc=0.2w_c = 0.2.

Candidate B, in full:

UB=0.3(0.50)+0.5(0.80)+0.2(0.57)=0.15+0.40+0.114=0.664U_B = 0.3(0.50) + 0.5(0.80) + 0.2(0.57) = 0.15 + 0.40 + 0.114 = 0.664

The other two:

  • UA=0.3(1.00)+0.5(0.00)+0.2(0.00)=0.300U_A = 0.3(1.00) + 0.5(0.00) + 0.2(0.00) = 0.300
  • UC=0.3(0.00)+0.5(1.00)+0.2(1.00)=0.700U_C = 0.3(0.00) + 0.5(1.00) + 0.2(1.00) = 0.700

Ranking: C (0.700) > B (0.664) > A (0.300). Candidate C wins, but only barely over B — and it wins on latency and cost, not quality. Candidate A, the quality leader, finishes last because it is the worst on both other metrics and the weights don't reward quality enough to compensate.

Knowledge check

Check your understanding

Answer this question before you continue.

In the worked example, candidate B has normalized utilities 0.50 for quality, 0.80 for latency, and 0.57 for cost. With weights 0.3, 0.5, and 0.2 respectively, what is its weighted score?
Output Prediction

Focus: Compute a weighted candidate score from normalized utilities and a stated weight vector.

Change One Weight, Watch the Ranking Move

Hold the raw values and bounds fixed. Change only the weights to reflect a cost-sensitive stakeholder: wq=0.3w_q = 0.3, wl=0.2w_l = 0.2, wc=0.5w_c = 0.5.

  • UA=0.3(1.00)+0.2(0.00)+0.5(0.00)=0.300U_A = 0.3(1.00) + 0.2(0.00) + 0.5(0.00) = 0.300
  • UB=0.3(0.50)+0.2(0.80)+0.5(0.57)=0.15+0.16+0.285=0.595U_B = 0.3(0.50) + 0.2(0.80) + 0.5(0.57) = 0.15 + 0.16 + 0.285 = 0.595
  • UC=0.3(0.00)+0.2(1.00)+0.5(1.00)=0.700U_C = 0.3(0.00) + 0.2(1.00) + 0.5(1.00) = 0.700

Ranking: C (0.700) > B (0.595) > A (0.300). The order holds, but the gap between C and B widens from 0.036 to 0.105. C is now more clearly the choice.

Push further. What if quality matters most — wq=0.6w_q = 0.6, wl=0.2w_l = 0.2, wc=0.2w_c = 0.2?

  • UA=0.6(1.00)+0.2(0.00)+0.2(0.00)=0.600U_A = 0.6(1.00) + 0.2(0.00) + 0.2(0.00) = 0.600
  • UB=0.6(0.50)+0.2(0.80)+0.2(0.57)=0.30+0.16+0.114=0.574U_B = 0.6(0.50) + 0.2(0.80) + 0.2(0.57) = 0.30 + 0.16 + 0.114 = 0.574
  • UC=0.6(0.00)+0.2(1.00)+0.2(1.00)=0.400U_C = 0.6(0.00) + 0.2(1.00) + 0.2(1.00) = 0.400

Ranking: A (0.600) > B (0.574) > C (0.400). The winner flips. Candidate A, dead last under the latency-sensitive weights, takes first place once quality dominates.

The mechanism is simple: the ranking is a function of the weights. A ranking without stated weights is an incomplete claim. Report the weight vector alongside the winner, and test at least two plausible weight vectors before committing.

Knowledge check

Check your understanding

Answer this question before you continue.

With the worked-example bounds, the weights change to quality 0.6, latency 0.2, and cost 0.2. Which candidate ranks first, and why?
Comparison Reasoning

Focus: Determine how changing stakeholder weights can change the winning candidate while measurements and bounds stay fixed.

Change the Bounds, Watch the Ranking Move Again

Bounds get less attention than weights and do just as much work. This time, switch to a fixed decision-range bound: suppose your product team has decided that quality 10 is the best value worth caring about, not quality 8. The candidates are unchanged. Only the anchor moved.

Keep the latency-sensitive weights (wq=0.3w_q = 0.3, wl=0.5w_l = 0.5, wc=0.2w_c = 0.2) and set qmin⁡=6q_{\min}=6, qmax⁡=10q_{\max}=10.

New quality utilities: uq(A)=(8−6)/4=0.50u_q(A) = (8-6)/4 = 0.50, uq(B)=(7−6)/4=0.25u_q(B) = (7-6)/4 = 0.25, uq(C)=0u_q(C) = 0.

  • UA=0.3(0.50)+0.5(0.00)+0.2(0.00)=0.150U_A = 0.3(0.50) + 0.5(0.00) + 0.2(0.00) = 0.150
  • UB=0.3(0.25)+0.5(0.80)+0.2(0.57)=0.075+0.40+0.114=0.589U_B = 0.3(0.25) + 0.5(0.80) + 0.2(0.57) = 0.075 + 0.40 + 0.114 = 0.589
  • UC=0.3(0.00)+0.5(1.00)+0.2(1.00)=0.700U_C = 0.3(0.00) + 0.5(1.00) + 0.2(1.00) = 0.700

The order holds, but every quality utility compressed. A wider range stretches the denominator and squeezes candidates toward the middle. A narrower range does the opposite.

Now compare the two bound policies directly. Under candidate-derived bounds, A scored 1.00 on quality because it was the best in the set. Under fixed bounds anchored at 10, A scores 0.50 because it only reaches half the range you care about. Same candidate, same measurement, different frame — and a different score.

This is the quiet hazard. Under candidate-derived bounds, adding a very strong or very weak candidate to the comparison set silently changes every other candidate's score. No measurement changed. Only the frame did. Under fixed bounds, the scores stay stable, but you have to defend the anchor itself: why is 10 the ceiling and not 8 or 12?

Common mistake: Mixing the two policies. If you derive bounds from the candidate set in one analysis and from external targets in the next, your scores are not comparable. Pick one policy, name it in your report, and stay consistent.

What the Single Scalar Hides

The score is a preference aggregator, not a decision-maker. Four things it will happily hide from you:

Hard constraints are not preferences. A latency ceiling or a compliance requirement is a gate, not a weight. Filter candidates first, then score the survivors. A candidate that violates a constraint should never enter the arithmetic.

Distinct objectives can be non-substitutable. If two metrics must both clear a bar, an additive score will trade one away without complaint. The linear model cannot represent "both, or neither."

Quality is the softest input. If qq comes from a benchmark that does not match your workload, the score is precise about the wrong thing. Precision is not accuracy.

Averaging hides variance. Two candidates with the same UU can behave very differently under load. The score tells you the expected case, not the worst case.

A short checklist before you present a ranking:

  1. Filter on constraints.
  2. Score on preferences.
  3. Report weights, bounds, and the bound policy you used.
  4. Stress-test at least one weight and one bound.

The Decision Rule

Never present a ranking without its weights and its bounds. Those two choices are the argument. The arithmetic is just the receipt.

Take your own two or three real candidate configurations. Write down the bounds, the bound policy, and the weights before you score anything — that ordering is what keeps the numbers honest. Then run the sensitivity test on the weight you are least sure about, and watch whether the winner survives.

The point is not to end the disagreement. It is to make the disagreement visible, specific, and testable. "I think latency should carry more weight" is a claim you can now settle with a re-run instead of a re-meeting.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

A product has a strict latency ceiling, and one candidate violates it. According to the article's decision process, what should the team do before comparing weighted utility scores?
Question 1 of 2Scenario Interpretation

Focus: Distinguish hard constraints from preferences in a weighted-utility decision process.

Two candidates receive the same weighted score, but their behavior under load varies differently. What conclusion is supported by the article?
Question 2 of 2Misconception Check

Focus: Explain a limitation of interpreting a weighted scalar as a complete account of candidate behavior.

Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.