Skip to content
intermediate

How Evaluation Population Weights Change an LLM's Overall Score

Two teams run the same evaluation suite on the same two systems. One publishes a table where System A wins. The other publishes a table where System B…

Published 2026-10-03Updated 2026-10-0411 min read
Vast desert landscape featuring rolling sand dunes under a clear sky.
Vast desert landscape featuring rolling sand dunes under a clear sky. Photo by Christophe RASCLE on Pexels.

Two teams run the same evaluation suite on the same two systems. One publishes a table where System A wins. The other publishes a table where System B wins. No model changed. No prompt changed. No test case changed.

Only the assumed traffic mix changed.

That should bother you, because it means an aggregate benchmark score is not a property of a model. It is a property of a model under a stated population assumption — and that assumption is usually buried inside the aggregation rule, unprinted, unexamined, and quietly doing most of the work.

By the end of this article you will be able to write down the strata, the stratum scores, and the weights behind any aggregate number, recompute it under a different deployment mix, and watch a ranking flip on paper.

Why Two Teams Get Opposite Rankings From the Same Results

The default belief is simple and mostly wrong: an overall score is a property of the model, so a higher score means a better model. Under that belief, comparing two systems is just reading two numbers.

There is a narrow case where that belief holds. If the evaluation task mix genuinely matches your deployment population, and that mix is stable over time, then the aggregate is a reasonable summary of expected performance for your traffic. The belief is not stupid. It is just conditional, and the condition is rarely checked.

The hidden constraint is that the weights are baked into the aggregation rule. When someone averages 500 test cases, they are implicitly asserting that every case matters equally often. When someone reports a headline number from a benchmark, they are implicitly asserting that the benchmark's task distribution is the one you care about. Neither assertion is printed next to the score. The score arrives clean, decimal and confident, with its assumptions stripped off.

So when two teams disagree, the disagreement usually is not about the models. It is about which population each team silently assumed.

Task Strata, Stratum Scores, and Population Weights

Let's make the hidden machinery visible. Three objects do all the work.

A task stratum is a group of evaluation cases that share a task type, risk level, or user intent. Short factual lookups are a stratum. Multi-step reasoning is a stratum. Safety-sensitive refusals are a stratum. A stratum is not an arbitrary bucket of leftover cases — it is a category you would defend as meaningfully different in behavior.

The stratum score sis_i is the mean score of the cases inside stratum ii. Two things matter here. First, the within-stratum aggregation rule is itself a choice: mean, pass rate, or geometric mean will give you different numbers. Second, sis_i is only meaningful if the cases inside the stratum are comparable on the same scale.

The population weight wiw_i is the assumed share of real deployment traffic that falls into stratum ii. Weights are non-negative and sum to 1:

∑iwi=1\sum_i w_i = 1

The weighted score is then:

S=∑iwi⋅siS = \sum_i w_i \cdot s_i

Read that aloud in plain causal language: each stratum contributes to the total in proportion to how often you expect to meet it. A stratum you meet constantly pulls the score hard. A stratum you meet once a month barely moves it — no matter how badly it fails.

Note: The formula carries an assumption stack. Strata are disjoint (no case belongs to two), weights are fixed and known, and stratum scores are comparable on the same scale. Break any of those and the weighted sum stops meaning what you think it means.

Knowledge check

Check your understanding

Answer this question before you continue.

An evaluation groups cases by task type. What does the population weight for one stratum represent?
Single Choice

Focus: Distinguish deployment-population weights from within-stratum scores.

Deriving the Weighted Score Step by Step

Start with the unweighted mean over all cases. If you have NN cases and case jj scores xjx_j, the mean is:

xˉ=1N∑j=1Nxj\bar{x} = \frac{1}{N} \sum_{j=1}^{N} x_j

Now group the cases by stratum. Let stratum ii contain nin_i cases, so ∑ini=N\sum_i n_i = N. The same sum, reorganized:

xˉ=1N∑i∑j∈ixj=∑iniN⋅1ni∑j∈ixj\bar{x} = \frac{1}{N} \sum_i \sum_{j \in i} x_j = \sum_i \frac{n_i}{N} \cdot \frac{1}{n_i}\sum_{j \in i} x_j

The inner term is exactly sis_i, the stratum score. The outer factor ni/Nn_i / N is the stratum's share of case count. So:

xˉ=∑i(niN)⋅si\bar{x} = \sum_i \left(\frac{n_i}{N}\right) \cdot s_i

The unweighted mean is a weighted sum in disguise, with weights equal to each stratum's share of the test set. This is the punchline: the uniform-weight case is a special case, not a neutral default. Averaging every case equally silently asserts that every stratum matters equally often. If your real traffic is 80% short lookups and your test set is 33% short lookups, the plain mean is not neutral — it is wrong in a specific, directional way.

The general weighted form replaces ni/Nn_i / N with wiw_i, the deployment-population share. Same algebra, different assumption.

What breaks the whole construction? Overlapping strata double-count cases, inflating whichever stratum sits in the overlap. Non-comparable stratum scales — one scored 0–1, another 1–10 — make the weighted sum a category error, not a number. And a stratum with very few cases produces an unstable sis_i that the weight then either amplifies into the headline or buries.

Knowledge check

Check your understanding

Answer this question before you continue.

A team averages every test case equally, even though its task strata contain different numbers of cases. Which interpretation of that overall mean is correct?
Misconception Check

Focus: Explain which weights are implicit in an unweighted mean over all evaluation cases.

Worked Example: One Score, Two Deployment Mixes

A comparison matrix lists fixed scores for Systems A and B across three task strata, alongside two different weight mixes. The totals show A ahead under Mix A and B ahead under Mix B.
The model scores stay the same; changing the assumed task mix flips the aggregate winner.

Three strata, two systems, two assumed populations. Every stratum score stays fixed throughout. Only the weights move.

StratumSystem ASystem BMix A weightMix B weight
Short factual lookups0.920.800.700.20
Multi-step reasoning0.600.850.200.50
Safety-sensitive refusals0.950.900.100.30

Mix A is a consumer chat profile: mostly quick lookups, occasional reasoning, rare safety-sensitive turns. Mix B is an internal analyst profile: heavy reasoning, more safety-sensitive content, fewer trivial lookups.

Under Mix A:

SA=0.70(0.92)+0.20(0.60)+0.10(0.95)=0.644+0.120+0.095=0.859S_A = 0.70(0.92) + 0.20(0.60) + 0.10(0.95) = 0.644 + 0.120 + 0.095 = 0.859

SB=0.70(0.80)+0.20(0.85)+0.10(0.90)=0.560+0.170+0.090=0.820S_B = 0.70(0.80) + 0.20(0.85) + 0.10(0.90) = 0.560 + 0.170 + 0.090 = 0.820

System A wins by roughly four points. The lookup stratum dominates, and A is much better at lookups.

Now recompute under Mix B:

SA=0.20(0.92)+0.50(0.60)+0.30(0.95)=0.184+0.300+0.285=0.769S_A = 0.20(0.92) + 0.50(0.60) + 0.30(0.95) = 0.184 + 0.300 + 0.285 = 0.769

SB=0.20(0.80)+0.50(0.85)+0.30(0.90)=0.160+0.425+0.270=0.855S_B = 0.20(0.80) + 0.50(0.85) + 0.30(0.90) = 0.160 + 0.425 + 0.270 = 0.855

System B wins by nearly nine points. Same models. Same stratum scores. The ranking flipped because the reasoning stratum went from 20% to 50% of the assumed traffic, and B is much better at reasoning.

Look at the contribution table to see which stratum is doing the work:

StratumMix A contribution (A / B)Mix B contribution (A / B)
Short factual lookups0.644 / 0.5600.184 / 0.160
Multi-step reasoning0.120 / 0.1700.300 / 0.425
Safety-sensitive refusals0.095 / 0.0900.285 / 0.270

Under Mix A, the lookup stratum contributes 0.644 of A's 0.859 — about 75% of the total. Under Mix B, reasoning contributes 0.300 of A's 0.769 and 0.425 of B's 0.855. The flip lives entirely in the weights.

Tip: When a ranking flips between two reports, do not re-run the models. Diff the weights first. The flip is almost always there.

Knowledge check

Check your understanding

Answer this question before you continue.

Using Mix B, what weighted score does System B receive?
Output Prediction

Focus: Calculate an aggregate score from fixed stratum scores and a stated deployment mix.

Mix B weights: short factual lookups 0.20, multi-step reasoning 0.50, safety-sensitive refusals 0.30.
System B stratum scores: 0.80, 0.85, 0.90.

How an Overall Score Conceals Subgroup Behavior

The arithmetic above is not just a curiosity. It is a concealment mechanism, and it runs in both directions.

A high aggregate can hide a stratum where the system fails badly, as long as that stratum is rare under the assumed mix. System A scores 0.859 under Mix A while sitting at 0.60 on reasoning. If your users occasionally ask multi-step questions, that 0.60 is a real liability the headline number never mentions.

A low aggregate can hide a stratum where the system is excellent and which dominates your actual traffic. A system that scores poorly overall might be the best choice for a narrow, high-volume workload — if you re-weight to your own mix.

Weight sensitivity is the first failure mode. A point estimate is not enough. Report how much the score moves when weights shift within a plausible range. If a 10% shift in the reasoning weight moves the ranking, your conclusion is fragile and you should say so.

Small-stratum noise is the second. A stratum with few cases produces an unstable sis_i. A stratum with 12 cases and a 0.95 score is not evidence of reliability; it is a coin that landed heads twelve times.

Common mistake: Publishing a weighted score without the weights, the stratum scores, and the case counts beside it. A number without its assumptions is not evidence. It is a rumor with a decimal point.

Knowledge check

Check your understanding

Answer this question before you continue.

System A scores 0.859 under Mix A but only 0.60 on multi-step reasoning, a stratum weighted at 0.20. What is the most appropriate interpretation?
Scenario Interpretation

Focus: Explain how a high aggregate can conceal weak performance in a low-weight stratum.

When Weighted Aggregation Helps and When It Misleads

Weighted aggregation is the right tool when three conditions hold: you have a defensible traffic estimate, your strata are stable over time, and the decision genuinely depends on expected-population performance. Under those conditions, the weighted score is the honest summary.

It misleads when you use it to compare systems across different populations. A score computed under Mix A and a score computed under Mix B are not comparable numbers, even if they look identical on a slide. It misleads when the weights are guesses dressed as measurements — a weight you cannot defend is an opinion with decimal places.

Now the boundary that trips people up. Not every safety-related number belongs in the weighted sum. There are two distinct things hiding under the word "safety," and they need opposite treatment.

Safety-related quality is performance on tasks that involve sensitive content but still have a graded quality dimension — how well a system handles a delicate medical question, how clearly it explains a risky procedure, how gracefully it redirects. That kind of performance is substitutable with other quality dimensions and can legitimately enter a weighted score. It is what the "safety-sensitive refusals" stratum in the example above represents: a quality score on sensitive turns, averaged like any other stratum.

A minimum safety requirement is different. It is a non-compensable gate: a specific policy violation, a prohibited output, a failure that must not be traded against arithmetic skill or latency. A gate does not get a weight. It gets a pass/fail check that runs before the weighted score is even computed. If the gate fails, the aggregate is irrelevant.

Warning: Keep minimum safety requirements as separate gates, not as weighted components. A system that trips a policy gate should fail outright, regardless of how well it scores elsewhere. Averaging is for substitutable quality; it is not for compliance.

So the example's refusal stratum is fine as a weighted component because it measures graded quality on sensitive turns. If you redefined that stratum to mean "did the system ever emit a prohibited output," it would stop being a weighted component and become a gate. Same label, different object, different treatment. Decide which one you are measuring before you decide where it goes.

The practical alternative I would ship: report a small vector of stratum scores plus one clearly labeled weighted summary, and let the reader re-weight it. That gives you the compression of a single number without hiding the structure underneath it.

Where the Weights Come From in Practice

Weights come from four places, and they carry different evidence strength. Production traffic logs are the strongest — they are observations. Product requirements are next — they are commitments. Stakeholder judgment is weaker — it is a belief. A deliberate uniform baseline is weakest of all, but it is at least honest about being a baseline.

State the provenance of the weights next to the score. "Weights from Q3 production traffic, n=48,000 requests" is a defensible line. "Weights from team discussion" is also fine, as long as you say so.

Revisit the weights when the product, user base, or feature mix changes. A stale mix silently invalidates every comparison built on it, and nothing in the report will warn you.

This article sits downstream of the case-construction work: the strata you weight are only as good as the representative cases inside them. It sits upstream of regression comparison, where you re-run the same weighted score across versions to catch quality drops. And it depends on keeping retrieval failures separate from generation failures, because a stratum score that mixes both will hide which layer actually broke.

The Next Move

Before you trust any aggregate score — yours or someone else's — ask three questions. What population does it assume? Is that population mine? What does the score hide?

Then do one concrete thing. Take an existing evaluation result, write down the stratum scores and the weights it implies, and recompute it under your own traffic mix. If the conclusion survives the re-weighting, you have evidence. If it flips, you just found the assumption that was making the decision for you.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

A team is deciding whether two safety-related measures belong in a weighted score. Which treatment matches the article's distinction?
Question 1 of 2Comparison Reasoning

Focus: Differentiate graded safety-related quality from a non-compensable minimum safety requirement.

A team wants weights that reflect its current deployment traffic. Which approach provides the strongest evidence and makes the assumption auditable?
Question 2 of 2Scenario Interpretation

Focus: Identify the strongest evidence source for deployment-population weights and the need to disclose its provenance.

Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.