Skip to content
intermediate

Practice Finding a Pareto Frontier for LLM Tradeoffs

Six configurations sit on the whiteboard. Someone reaches for a single score to rank them, and the interesting part of the decision disappears.

Published 2026-10-03Updated 2026-10-0410 min read
A beautiful close-up of sand dunes in a desert landscape, showcasing natural patterns.
A beautiful close-up of sand dunes in a desert landscape, showcasing natural patterns. Photo by MART PRODUCTION on Pexels.

Six configurations sit on the whiteboard. Someone reaches for a single score to rank them, and the interesting part of the decision disappears.

Here is the move I want you to practice instead: before you weight anything, find the frontier. The frontier is the set of options that are still genuinely different bets. Everything else is a configuration you can delete without losing a single tradeoff.

This is a hands-on exercise. You will work one small synthetic table, mark the dominated rows, read the surviving frontier, then change one constraint and watch the answer move. The numbers are invented to produce a useful mix of dominated and non-dominated points. They do not describe any real provider, and they are not a recommendation.

If you have not yet read the prerequisite piece on cost, latency, and quality tradeoffs, the short version is enough: those three objectives pull against each other, and you cannot maximize all of them at once. That article explains why the tension exists. This one is about executing the comparison.

The Fixture: Seven Configurations, Three Objectives

ConfigQuality score (higher is better)p95 latency, ms (lower is better)Cost per 1k requests, $ (lower is better)
A719004.20
B7814006.50
C7411005.10
D8221009.80
E698003.40
F7613007.90
G7312006.80

Three columns, three directions. Quality goes up. Latency goes down. Cost goes down.

Common mistake: Flipping a metric direction changes every dominance verdict in the table. If you treat "lower latency" as "higher is better" without inverting it, you will confidently eliminate the wrong rows. Write the direction next to each column before you compare anything.

The numbers are synthetic and deliberately small. Seven rows is enough to practice the procedure and small enough that you can check every pair by hand. This table stays fixed for the rest of the exercise — every result below is computed from exactly these values.

What Dominance Actually Means

A configuration is dominated when another configuration is at least as good on every objective and strictly better on at least one.

That is the whole rule. Two conditions, both required.

Walk one pair by hand. Compare A and C:

  • Quality: C (74) beats A (71).
  • Latency: A (900 ms) beats C (1100 ms).
  • Cost: A ($4.20) beats C ($5.10).

C wins one column, A wins two. Neither dominates the other. Both stay live for now.

Now compare A and E:

  • Quality: A (71) beats E (69).
  • Latency: E (800 ms) beats A (900 ms).
  • Cost: E ($3.40) beats A ($4.20).

Again, a split. A wins quality, E wins latency and cost. No domination.

Here is the trap. A configuration that wins on quality but loses on both latency and cost is not dominated. It is a different bet. It stays on the table until a constraint or a weight removes it. Beginners often delete the expensive high-quality option because it "feels worse." It is not worse. It is the only row that buys you that quality level.

Knowledge check

Check your understanding

Answer this question before you continue.

Configuration X is better than Y on quality but worse on latency and cost. What follows under the article's dominance rule?
Single Choice

Focus: Apply the strict dominance rule to distinguish dominated configurations from genuine tradeoffs.

Step 1: Mark the Dominated Rows

Run this routine for each configuration:

  1. Pick a row.
  2. Scan every other row.
  3. Ask: does any single row beat this one on all three metrics, with at least one strict win?
  4. If yes, mark the row dominated and record which row beat it and on what.

Work it through. Start with B (78, 1400, $6.50). Scan the table. D has higher quality (82) but worse latency and cost. A is cheaper and faster but lower quality. C is cheaper and faster but lower quality. No single row beats B on all three. B survives.

Now C (74, 1100, $5.10). Compare against B: B has higher quality (78), but worse latency (1400 vs 1100) and worse cost ($6.50 vs $5.10). Split. Compare against A: A is faster and cheaper but lower quality. Split. C survives.

Now F (76, 1300, $7.90). Compare against B:

  • Quality: B (78) beats F (76).
  • Latency: F (1300) beats B (1400).
  • Cost: B ($6.50) beats F ($7.90).

Split again. F survives.

Now D (82, 2100, $9.80). Highest quality in the table. Nothing beats it on quality, so nothing can dominate it. D survives by construction.

Now G (73, 1200, $6.80). Compare against C (74, 1100, $5.10):

  • Quality: C (74) beats G (73).
  • Latency: C (1100) beats G (1200).
  • Cost: C ($5.10) beats G ($6.80).

C beats G on all three, with strict wins everywhere. C dominates G. G is eliminated, and you can delete it from the whiteboard without losing a tradeoff.

That leaves A and E. Compare A (71, 900, $4.20) against E (69, 800, $3.40):

  • Quality: A (71) beats E (69).
  • Latency: E (800) beats A (900).
  • Cost: E ($3.40) beats A ($4.20).

Split. A survives. E survives by the same logic in reverse.

Note: A near-miss deserves a caveat. If two rows differ by a hair on one metric — say 1100 ms versus 1120 ms — do not declare a confident verdict. Small differences can be measurement noise. Note the pair and move on.

Knowledge check

Check your understanding

Answer this question before you continue.

Which comparison establishes that G is dominated in the fixture?
Comparison Reasoning

Focus: Use the stated metric directions to identify a configuration that is strictly dominated in the fixture.

Step 2: Read the Frontier

A scatter plot places quality on the vertical axis and cost on the horizontal axis. Configurations E, A, C, F, B, and D are highlighted as frontier members; G is marked as dominated. Each point is labeled with its configuration and latency in milliseconds, showing why F remains on the frontier despite its higher cost than B: F has lower latency.
The plot makes the surviving tradeoffs visible; latency labels preserve the third objective that a two-axis view would otherwise hide.

The surviving set is A, B, C, D, E, and F. Six non-dominated configurations. Read them left to right by quality:

ConfigQualityp95 latency, msCost per 1k, $
E698003.40
A719004.20
C7411005.10
F7613007.90
B7814006.50
D8221009.80

The broad shape is what you would expect: as quality rises, latency and cost tend to rise with it. But the marginal trade is not uniform, and one step breaks the pattern. Walk the steps:

  • E to A: +2 quality for +100 ms and +$0.80.
  • A to C: +3 quality for +200 ms and +$0.90.
  • C to F: +2 quality for +200 ms and +$2.80.
  • F to B: +2 quality for +100 ms and −$1.40 (B is cheaper than F).
  • B to D: +4 quality for +700 ms and +$3.30.

That F-to-B step is the interesting one. B is both higher quality and cheaper than F, but slower. F is only on the frontier because of its latency. If your application does not care about the 100 ms difference, F is a poor deal — but it is not dominated, because latency is a real objective and F genuinely wins that column.

A two-axis view helps here: plot quality on the vertical axis and cost on the horizontal, then annotate each point with its latency. The frontier traces a curve from bottom-left (E) to top-right (D), with F sitting slightly off the cost curve because its value is speed, not price.

The key point: the frontier does not rank its own members. It tells you which choices are still live. It does not tell you which one to pick.

Knowledge check

Check your understanding

Answer this question before you continue.

Using higher quality and lower latency and cost as better, which set is the fixture's frontier?
Output Prediction

Focus: Predict the non-dominated set for the supplied seven-configuration fixture.

Step 3: Change One Constraint and Recompute

Feasibility is a separate question from dominance. Dominance is a property of the metrics. Feasibility is a property of your requirements.

Suppose your product has a hard p95 latency ceiling of 1200 ms. Apply it to the frontier:

ConfigLatencyFeasible under 1200 ms?
E800Yes
A900Yes
C1100Yes
F1300No
B1400No
D2100No

The frontier shrinks from six members to three. F, B, and D are now infeasible — not dominated, just unavailable. Their metric values did not change. Your requirement did.

Now check whether any previously dominated row becomes relevant. G was dominated by C, and C is still feasible, so G stays eliminated. But imagine the constraint had been a cost cap of $5.00 instead. Then C ($5.10) would be infeasible, and G ($6.80) would still be infeasible — but a different row might resurface. The lesson holds: when a constraint removes a dominator, re-run the dominance check on what remains.

Warning: If the constraint eliminates the entire frontier, your requirement set is unsatisfiable. That is a real signal, not a bug. Something has to give: relax the latency ceiling, accept higher cost, or lower the quality bar. Do not pretend a dominated row is now optimal just because it is the only one left.

Knowledge check

Check your understanding

Answer this question before you continue.

If the hard p95 latency ceiling is 1200 ms, which configurations remain feasible members of the original frontier?
Scenario Interpretation

Focus: Recompute the feasible portion of a frontier after imposing a hard latency ceiling.

Why This Frontier Is Not a Recommendation

The frontier you just computed is conditional on four things: the metric directions, the metrics you chose, how you defined them, and the constraint you applied. Change any one and the answer moves.

Latency is the clearest example. A p95 time-to-first-token measurement and a p95 total-latency measurement produce different frontiers from the same system. If your users judge responsiveness by when the first token appears, TTFT is the metric that matters. If they wait for the full answer, total latency is. Pick the wrong one and you optimize the wrong curve.

The fixture is also synthetic and small. Real comparisons need consistent measurement conditions — same hardware, same load, same prompt distribution — and an awareness that small metric gaps can be noise rather than signal.

The durable rule: the frontier narrows the decision; it does not make it. Weighting and constraints still belong to the person with the requirement. That is the next step, and it is a separate exercise.

Your Turn: Extend the Fixture

Add a seventh configuration, H, with quality 80, latency 3000 ms, and cost $4.00. Predict where it lands before you check.

Work the comparisons explicitly. Against D (82, 2100, $9.80): D wins quality, H wins latency and cost — split, so D does not dominate H. Against B (78, 1400, $6.50): B wins latency and quality, H wins cost — split, so B does not dominate H. Against E (69, 800, $3.40): E wins latency and cost, H wins quality — split, so E does not dominate H. No single row beats H on all three metrics, so H is non-dominated. It is a genuine different bet: cheap and high-quality, but the slowest option in the table.

Then add a fourth objective, such as output token variance or memory footprint, and re-run the dominance check. Watch how quickly the frontier changes shape when you add a dimension. More objectives usually mean more surviving rows, because it becomes harder for any single configuration to beat another on every column.

Finally, flip one metric direction by mistake — treat cost as "higher is better" — and re-run Step 1. Feel how fast a careless sign error rewrites the answer. That is the failure mode this exercise exists to prevent.

Once the frontier is clear, the weighted-utility approach is the natural next move: assign weights to quality, latency, and cost, normalize each metric, and compute a single score. But do that after you know which options are live. Find the frontier first. Apply constraints second. Weight last.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

What does the computed frontier establish, and what does it leave to the decision-maker?
Question 1 of 2Misconception Check

Focus: Explain why a computed frontier is conditional rather than a universal model recommendation.

The exercise adds H with quality 80, latency 3000 ms, and cost $4.00. What is its status when compared with the existing configurations?
Question 2 of 2Output Prediction

Focus: Evaluate a new configuration against the fixture by checking whether any single existing option dominates it on every objective.

References

  1. Economic Evaluation of LLMsarxiv.org
  2. One Benchmark Is Not Enough to Choose Your Inference Providerfriendli.ai
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.