Skip to content
advanced

LLM Evaluation Validity: What Do Your Labels Actually Measure?

Two teams evaluate the same model on the same task and report 82% and 71%. Neither team is lying. Neither team is incompetent. They are running two…

Published 2026-10-03Updated 2026-10-0411 min read
Vast desert with rolling sand dunes under a clear blue sky, epitomizing isolation and natural beauty.
Vast desert with rolling sand dunes under a clear blue sky, epitomizing isolation and natural beauty. Photo by Matheus Natan on Pexels.

Two teams evaluate the same model on the same task and report 82% and 71%. Neither team is lying. Neither team is incompetent. They are running two different measurement instruments and calling both of them "quality."

That gap is not sampling noise, and it is not a judge that needs recalibrating. It is a validity problem, and it lives upstream of almost everything else you do with an eval set. Before you can trust a score, you have to know what the label underneath it actually measures.

The Construct, the Observation, and the Label

A four-stage evaluation path moves from the intended construct, such as a reply being sendable without edits, to the observed response, through a rubric-based decision rule, and to a pass or fail label. A separate real-world outcome connects back to the label as an external validity check.
A label is produced from observations by a decision rule; check whether it tracks the consequence the construct is meant to represent.

Most evaluation discussions collapse three distinct layers into one word: "quality." Separate them and the failure modes become visible.

The latent construct is the thing you actually care about. Not "the answer is good," but something like "a support agent can send this reply without editing it." The construct is latent because it exists in the world of consequences, not in the text.

The observation is the evidence you can see: the response text, the retrieved context, the tool trace, the latency, the token count. Observations are cheap and plentiful. Constructs are expensive and rare.

The operational label is the discrete value a human annotator or automated grader assigns to an observation. Pass. Fail. 1 through 5. This is what ends up in your spreadsheet and your dashboard.

The mapping between observation and label is a decision rule, and it is worth writing down explicitly:

label = f(observation, rubric, annotator, threshold)

Every term in that function is a place where validity can break. The rubric can be ambiguous. The annotator can hold a different internal construct than the one you intended. The threshold can be arbitrary. Change any one of them and the label moves, even when the observation is identical.

Here is the assumption that makes the whole chain work: the label is a monotone, low-noise function of the construct. Higher construct quality should produce higher label values, reliably. That assumption is almost never tested. We write a rubric, label a few hundred examples, and start quoting percentages as if the mapping were a measurement instead of a decision.

A related trap: collapsing distinct constructs into one label. "Faithfulness" (does the answer stay grounded in the retrieved evidence?) and "helpfulness" (does the answer actually solve the user's problem?) are different constructs. A single pass/fail label that mixes them silently changes what your number means depending on which failure mode happened to dominate your sample.

Note: A label is not a measurement of quality. It is a decision procedure applied to an observation. Validity asks whether that procedure tracks the construct you claim to measure.

Knowledge check

Check your understanding

Answer this question before you continue.

For the billing-reply evaluation, which phrase names the latent construct rather than an observation or a label?
Single Choice

Focus: Distinguish a latent quality construct from observable evidence and the operational label assigned to that evidence.

A Bounded Annotation Example With Disagreement

Abstract vocabulary is easy to nod at. Let me make it concrete with a small, fully worked case.

Take a fixed set of 20 responses to one task type: a customer support assistant answering billing questions. The intended construct is "a support agent could send this reply with no edits." The label rule is binary: pass if the reply is factually correct and directly addresses the question; fail otherwise. The rubric is written down.

Two annotators, A and B, label all 20 items independently. They agree on 14. They disagree on 6.

Look at where they diverge. On four of the six disagreements, the reply is factually correct but hedged: "I believe the charge was applied on the 3rd, though you may want to confirm with your bank." Annotator A reads the hedge as appropriate caution and marks pass. Annotator B reads the hedge as a failure to commit and marks fail. Neither is violating the rubric as written, because the rubric never defined what "directly addresses" means for a hedged answer.

On the other two disagreements, the replies are genuinely borderline: partially correct, partially evasive. These are real edge cases, not rubric ambiguity.

Now compute the score three ways. The tally has to be auditable, so here is the full count. Annotator A passes 16 of 20 and fails 4. Annotator B passes 12 and fails 8. They agree on 14 items: 11 passes and 3 fails. That leaves 6 disagreements — the 4 hedge cases plus the 2 borderline cases. On 5 of those disagreements A says pass and B says fail; on 1, B says pass and A says fail. Check the totals: A's passes are 11 agreed + 5 A-only = 16. B's passes are 11 agreed + 1 B-only = 12. The arithmetic closes.

RuleScore
Annotator A only16 / 20 = 80%
Annotator B only12 / 20 = 60%
Majority (tie broken as fail)11 / 20 = 55%

Same system. Same observations. Three defensible numbers, spanning 25 points.

The disagreement is not one problem. It is at least two. The four hedge cases are rubric ambiguity — the rule underdetermines the label. The two borderline cases are genuine construct boundary — the annotators may hold slightly different internal constructs of "sendable." Each has a different repair. Ambiguity gets fixed by tightening the rubric. Boundary cases get fixed by deciding, explicitly, which side of the line you want to be on.

There is a third possibility worth naming: annotator drift, where one annotator's internal construct shifts over the course of a long labeling session. You cannot see it in a single pass, but you can see it by re-labeling a subset later and checking for self-disagreement.

Common mistake: Reporting raw agreement rate as evidence that labels are reliable. On an imbalanced set — say 90% pass — two annotators who both stamp "pass" on everything agree 90% of the time and have learned nothing. Per-class precision and recall expose what raw agreement hides.

Knowledge check

Check your understanding

Answer this question before you continue.

Using the article's 20-item tally and resolving the six disagreements by majority with ties counted as fail, what is the resulting pass score?
Output Prediction

Focus: Calculate the majority score from the article's annotator counts when ties are resolved as fail.

Change the Rule, Change the Conclusion

Here is the part that should make you nervous about every trend line you have ever drawn.

Hold the observations fixed. Change one thing: the treatment of hedged answers. Under the old rule, hedges count as pass. Under the new rule, hedges count as fail. Re-run the same comparison between two system versions — call them v1 and v2 — under both rules.

v1v2Verdict
Old rule (hedge = pass)74%78%v2 wins
New rule (hedge = fail)61%58%v1 wins

The ranking flips. The systems did not change. The instrument did.

This is the mechanism: a rule change is a change to the measurement instrument. Any before/after comparison that spans a rule change is comparing two instruments, not two systems. The number moved because the ruler moved.

The operational discipline is simple and almost never followed:

  • Freeze the label rule for the duration of a comparison. No mid-experiment rubric edits, no "we realized hedges should fail" after you have already collected half the labels.
  • Version the rule. Treat it like code. rubric_v3.md, dated, with a changelog.
  • Record the rule alongside the score. A score without its rule is a number with a story attached.

There is a legitimate case for changing the rule: when the construct genuinely changed. If your product pivoted from "draft replies for agents to edit" to "auto-send replies," then hedges really did become failures, and the rule should change. The distinction is whether the change is declared and versioned or silent drift. Silent drift invalidates every trend line that crosses it, and you will not notice until someone asks why the score dropped 15 points and no one can explain it.

Knowledge check

Check your understanding

Answer this question before you continue.

The comparison ranks v2 above v1 under the old hedge rule, but v1 above v2 under the new rule. What is the supported interpretation?
Comparison Reasoning

Focus: Explain why a changed operational label rule can reverse a system ranking without any change to the systems.

Validity Is Not Sampling Error and Not Judge Calibration

Three problems get conflated constantly. They have different symptoms, different diagnostics, and different fixes.

Sampling uncertainty is the problem of a finite set. Your 20 items are a sample; the true rate on the population has a confidence interval around your observed score. More items shrink the interval. More items do nothing for a label that measures the wrong construct. You can have a tight interval around a biased estimate.

Judge calibration is whether an automated grader agrees with your human labels. This matters — a judge that disagrees with humans is not useful as a proxy. But calibrating a judge against a biased label set produces a faithful copy of the bias. High agreement to a bad reference is not a virtue.

Measurement validity is whether the label tracks the construct at all. This is upstream of both. No amount of additional data and no amount of better agreement repairs a label that measures the wrong thing.

The diagnostic table below separates the two ideas that get merged most often: reliability (do annotators apply the rule consistently?) and validity (does the rule capture the construct?). Agreement checks tell you about reliability. They say nothing about whether the rule is the right rule.

ProblemLooks likeDiagnosed byFixed by
Sampling uncertaintyWide confidence intervalMore items, bootstrapCollect more data
Label reliabilityAnnotators apply the rule inconsistentlyTwo annotators, independent, on the same itemsTighten rubric, resolve boundary cases
Judge calibrationJudge disagrees with humansAgreement checks on held-out labelsBetter rubric, better judge prompt
Measurement validityScore moves for no system reason; labels miss the decision they were meant to serveChecking whether the rule predicts the downstream consequence (send-without-edit, ship/no-ship)Redefine construct or rule

Notice the last row. You cannot diagnose construct validity by counting agreements. Two annotators can agree perfectly on a rule that measures the wrong thing. The only real test is whether the label predicts the consequence the construct was supposed to represent — does "pass" actually mean an agent sent the reply unedited? That requires an external check, not an internal one.

The failure mode that should worry you most: a judge with high agreement to human labels that are themselves a poor proxy for the construct. The pipeline looks healthy. The agreement metric is green. The conclusion is still wrong, and nothing in your dashboard will tell you.

There is a related attribution problem. When a score drifts, the cause may be the system, the judge, or the label rule. These are distinguishable only if you keep a fixed, human-labeled anchor set that you re-score periodically. Without an anchor, every drift alarm is ambiguous between "the product got worse" and "the judge changed." A stable anchor set is the only thing that lets you tell them apart.

Knowledge check

Check your understanding

Answer this question before you continue.

A judge agrees closely with human labels, but the labels may not represent whether replies are actually sendable without edits. Which check most directly addresses measurement validity?
Misconception Check

Focus: Distinguish measurement validity from judge calibration and identify an external check of whether a label tracks its intended construct.

When This Analysis Is Worth the Cost

Construct-level validity work is not free. Here is where I would spend it and where I would not.

Worth it:

  • The label drives a ship/no-ship decision.
  • The construct is contested or multi-dimensional (faithfulness vs. helpfulness vs. tone).
  • The score is used to compare across time, across teams, or across model versions.

Overkill:

  • A narrow, mechanical construct with an objective check. Schema validity, exact match, unit tests — where the label is nearly the construct itself. If your check is "does this parse as valid JSON," you do not need two annotators and a disagreement analysis.

Cheap first moves that cost an afternoon and catch most of the damage:

  1. Write the decision rule down. Not the construct — the rule. What observation produces which label?
  2. Have two people label a small sample independently.
  3. Inspect the disagreements before building anything. The disagreements are the map of where your construct is underspecified.
  4. Check one external signal: does "pass" predict the real-world outcome you care about? If you can, sample a handful of passes and confirm they would have been sent unedited.

Warning: A single scalar score reported without its construct, its rule, or its disagreement rate is not evidence. It is a number with a story attached, and the story is doing the work the measurement should be doing.

The One-Sentence Test

Before you quote a score, be able to state four things in one sentence: the construct, the observation, the label rule, and the disagreement rate. If you cannot, you are reporting a measurement you have not validated — and the 82% and the 71% are both true, both defensible, and both measuring something you have not named.

The next practical step is small and unglamorous: write the decision rule down, then label a small sample twice. That is the foundation the rest of your evaluation workflow rests on. Regression tests, judge calibration, drift monitoring — all of it inherits whatever validity your labels have. Get the labels right first, or everything downstream is a more precise measurement of the wrong thing.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

A team uses a subjective, multi-dimensional quality score to make a ship/no-ship decision. Based on the article's cost guidance, what is the best next step?
Question 1 of 2Scenario Interpretation

Focus: Choose when construct-level validity analysis is worth its cost based on the stakes and nature of the construct.

A report names the intended construct, the response text being examined, and the pass/fail rule, but gives no disagreement rate. What does the article's one-sentence test imply?
Question 2 of 2Scenario Interpretation

Focus: Apply the article's one-sentence test by identifying the required elements needed to contextualize an evaluation score.

References

  1. [2504.19076] LLM-Evaluation Tropes: Perspectives on the Validity of LLM-Evaluationsarxiv.org
  2. Paper page - Who Drifted: the System or the Judge? Anytime-Valid Attribution in LLM Evaluation Pipelineshuggingface.co
  3. Demystifying evals for AI agents \ Anthropicwww.anthropic.com
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.