Practice Calibrating an LLM Judge Against Human Labels
Your judge reports a 92% pass rate. The dashboard is green. Then a reviewer reads ten of those passes and disagrees with four of them.

Key topics
Your judge reports a 92% pass rate. The dashboard is green. Then a reviewer reads ten of those passes and disagrees with four of them.
The judge did not malfunction. It produced numbers exactly as designed. The problem is that nobody checked whether those numbers mean what you think they mean. A judge is a measuring instrument, and an instrument that has never been calibrated will still give you a reading — it just won't be the reading you assume.
This exercise fixes that. You will take a small human-labeled set, score it with your judge, and then read the disagreements instead of averaging them away. By the end you will know where your judge is reliable, where it drifts, and whether it is fit for the specific decision you want it to support.
This assumes you have already decided model grading is appropriate for your task and have a rubric in hand. Here we only run the calibration loop.
What You Need Before You Start
Calibration fails at the setup stage far more often than at the analysis stage. Get these five things right first.
One narrow criterion. Pick a single thing to measure — "does the answer stay faithful to the provided context," not "is this answer good." A judge calibrated on everything is calibrated on nothing. If you need three criteria, run three separate calibration passes.
A frozen reference set of 30–60 real examples. Real production outputs, not synthetic ones. Include deliberately hard and borderline cases, because a set of easy examples will make any judge look aligned. Freeze it: once labeled, those items do not change, so a later score shift means the judge moved, not the data.
A written rubric with explicit level descriptions. Your human labels need a stable target. If reviewers cannot articulate why an answer earned its grade, the judge has nothing to match.
A pinned, dated judge model version and a stored judge prompt. Record both with every score. A model alias can be repointed under you, and a silent version change will quietly invalidate your calibration.
Human labels you trust enough to argue with. Before blaming the judge, check whether your reviewers agree with each other. If two humans disagree on a third of the set, the rubric is not specific enough yet — and that is the most valuable thing this exercise can tell you.
Common mistake: Labeling the set with a single reviewer and treating those labels as ground truth. Measure inter-rater agreement on a subset first. If your humans disagree, the judge is being graded against noise.
Run the Calibration Loop
Here is the smallest useful implementation. It takes a JSON file of human and judge labels, computes the confusion matrix, agreement statistics, and error direction, and prints a summary you can act on.
Environment: Python 3.9+ with no external dependencies. The script uses only the standard library.
Input format: A JSON file where each record has an item_id, a human_label ("pass" or "fail"), and a judge_label ("pass" or "fail"). Optionally include a judge_rationale string for later inspection.
[
{"item_id": "014", "human_label": "fail", "judge_label": "pass", "judge_rationale": "Answer restates the context accurately"},
{"item_id": "015", "human_label": "pass", "judge_label": "pass", "judge_rationale": "Grounded in the provided passage"},
{"item_id": "016", "human_label": "fail", "judge_label": "fail", "judge_rationale": "Adds a claim not present in context"}
]
import json
import sys
from collections import Counter
def load_labels(path):
with open(path) as f:
return json.load(f)
def confusion(records):
# Rows: judge label. Columns: human label.
matrix = {
("pass", "pass"): 0, ("pass", "fail"): 0,
("fail", "pass"): 0, ("fail", "fail"): 0,
}
for r in records:
matrix[(r["judge_label"], r["human_label"])] += 1
return matrix
def agreement(records):
n = len(records)
if n == 0:
return 0.0
agree = sum(1 for r in records if r["human_label"] == r["judge_label"])
return agree / n
def cohens_kappa(records):
n = len(records)
if n == 0:
return 0.0
human = Counter(r["human_label"] for r in records)
judge = Counter(r["judge_label"] for r in records)
p_observed = agreement(records)
p_expected = sum(
(human[label] / n) * (judge[label] / n)
for label in ("pass", "fail")
)
if p_expected == 1.0:
return 1.0
return (p_observed - p_expected) / (1.0 - p_expected)
def error_direction(matrix):
false_pass = matrix[("pass", "fail")]
false_fail = matrix[("fail", "pass")]
return false_pass, false_fail
def summarize(records):
m = confusion(records)
acc = agreement(records)
kappa = cohens_kappa(records)
fp, ff = error_direction(m)
print(f"Items: {len(records)}")
print(f"Raw agreement: {acc:.2%}")
print(f"Cohen's kappa: {kappa:.3f}")
print()
print("Confusion matrix (rows=judge, cols=human):")
print(f" judge=pass, human=pass: {m[('pass','pass')]}")
print(f" judge=pass, human=fail: {m[('pass','fail')]} <- false pass")
print(f" judge=fail, human=pass: {m[('fail','pass')]} <- false fail")
print(f" judge=fail, human=fail: {m[('fail','fail')]}")
print()
print(f"False passes: {fp} | False fails: {ff}")
if fp > ff:
print("Direction: judge is lenient (over-passes).")
elif ff > fp:
print("Direction: judge is strict (over-fails).")
else:
print("Direction: balanced error count.")
if __name__ == "__main__":
records = load_labels(sys.argv[1])
summarize(records)
Run it against your labeled file:
python calibrate.py labels.json
Expected output for a set where the judge over-passes:
Items: 40
Raw agreement: 82.50%
Cohen's kappa: 0.612
Confusion matrix (rows=judge, cols=human):
judge=pass, human=pass: 24
judge=pass, human=fail: 5 <- false pass
judge=fail, human=pass: 2 <- false fail
judge=fail, human=fail: 9
False passes: 5 | False fails: 2
Direction: judge is lenient (over-passes).
How to read this. Raw agreement of 82.5% sounds fine until you notice the judge is wrong in one direction five times more often than the other. Kappa of 0.612 tells you the agreement is real but not strong — well above chance, well below the level you would want for a release gate. The false-pass count is the number that should worry you if a false pass means a bad answer ships.
Do not compute a single accuracy number and stop. Build the confusion matrix first, because you need to see which direction the errors run. A judge that almost never says "fail" will look accurate on a set that is mostly "pass" — raw accuracy inflates when one label dominates. That is exactly why kappa is in the script: it corrects for agreement that would happen by luck. Kappa near zero means your judge is guessing with confidence.
Knowledge check
Check your understanding
Answer this question before you continue.
Read the Disagreements, Not the Average
The average tells you how often the judge agreed. The disagreements tell you what the judge is actually measuring. Sort every disagreement into one of three buckets:
- Rubric ambiguity — humans would also disagree here. The rubric needs a sharper boundary, not the judge.
- Judge error — the human label is clearly right and the judge is clearly wrong.
- Mishandled input class — the judge fails on a recognizable category: multi-part questions, answers that hedge, inputs with no relevant context.
Then read the judge's rationale on each disagreement. The stated reason is often more informative than the wrong label. A judge that says "the answer is detailed and well-structured" when the human marked it unfaithful is telling you it is scoring polish, not grounding.
Watch for proxies. A judge may be tracking something correlated with your criterion but not identical to it: answer length, confident tone, hedging, or formatting. If your criterion is faithfulness and your judge rewards length, you have built a verbosity detector with a faithfulness label on it.
Tip: Keep a running list of disagreement rationales. Those notes become the raw material for rubric fixes and anchor examples later — and they are the artifact your team will actually reuse.
Knowledge check
Check your understanding
Answer this question before you continue.
Run Three Bias Probes on the Same Set
Disagreements show you that the judge is off. Bias probes show you how it is systematically off. These are cheap experiments on the same frozen set.
Position and order probe. For pairwise or comparative judgments, evaluate both orderings of the same pair. If the verdict flips when you swap A and B, the judge has a position preference, and any comparison you run is partly a coin toss.
Verbosity probe. Correlate judge scores with answer length, then correlate human labels with answer length. If the judge tracks length harder than your humans do, it is measuring the wrong thing — and it will reward padding.
Self-preference probe. Check whether the judge scores outputs from its own model family higher than humans do. Where practical, use a judge from a different family than the generator.
Consistency probe. Rerun the judge on the same items and count verdict flips. Unstable items are your best candidates for human review, because the judge itself is uncertain.
The interpretation rule matters more than the probes: a bias that is small and stable may be tolerable for a bounded task. A bias that flips decisions at your threshold is not. Measure the bias, then ask whether it changes the decision.
Knowledge check
Check your understanding
Answer this question before you continue.
Decide Whether the Judge Is Good Enough
Now convert measurements into a bounded go / no-go decision. Start by stating the decision the judge will support — a release gate, a regression signal, or a triage filter — because the required agreement depends on the cost of being wrong.
Then compare agreement per criterion and per segment, not just overall. A judge can be dependable on easy cases and useless on the hard slice that actually matters. If your errors concentrate in one input class, fix the rubric or the judge prompt and re-run before declaring the judge unfit.
Set a threshold and a coverage tradeoff. You can demand higher agreement by letting the judge abstain on uncertain items and routing those to humans. That trades coverage for precision, and it is often the right trade for a narrow task.
Write down the failure direction you can tolerate. A false "pass" on a safety criterion is usually worse than a false "fail" on a style criterion. Your threshold should reflect that asymmetry, not a generic accuracy target.
| Decision type | What you need | What you can tolerate |
|---|---|---|
| Release gate | High agreement on the hard slice | Very few false passes |
| Regression signal | Stable agreement over time | Some noise, if it is consistent |
| Triage filter | High recall on failures | False alarms routed to humans |
Knowledge check
Check your understanding
Answer this question before you continue.
Failure Modes That Fool a Calibration Run
A calibration run that succeeds once can still produce false confidence. Watch for these.
Calibrating on easy or synthetic examples. The judge looks aligned because the set never tested the boundary. Sample real production outputs, weighted toward hard and high-stakes cases.
Rubric rot and drift. Agreement decays as traffic shifts, the model version changes, or the rubric stops matching what reviewers actually do. This is why you re-check on a schedule.
Treating human labels as ground truth when reviewers disagree. Measure inter-rater agreement before blaming the judge.
Overfitting the judge prompt to the reference set. Add anchor examples carefully. A prompt tuned until it memorizes those 40 items generalizes to nothing.
Changing the judge without a parallel run. Treat judge changes like a schema migration: run old and new judges side by side on the frozen set before cutting over.
Your Next Experiment: Tune on Development, Check on Holdout
Keep the frozen reference set frozen. Refresh a small sample of recent production traces on a cadence and label them against the same rubric. Track your agreement statistic over time as a leading indicator — a slow decline shows up here before it shows up in your quality dashboard.
For your next experiment, split your labeled items into two groups before you touch the judge prompt. Use one group — the development set — to pick the clearest disagreements and add them to the judge prompt as anchor examples. Keep the other group — the holdout set — untouched. Then run the revised judge on the holdout set and compare agreement before and after.
This matters because tuning a prompt on the same items you use to measure improvement is how you fool yourself. The judge can memorize those specific examples without getting any better at the underlying task. The holdout set is the only honest test of whether the change generalizes.
If agreement improves on the holdout, you have a real calibration win. If it improves only on the development set, you have overfitting. If neither improves, the problem is the rubric, not the prompt.
Decide in advance what would trigger a re-calibration versus a full judge rebuild. A judge is not trustworthy because it agrees on average. It is trustworthy because you know where it disagrees, in which direction, and on which slice. Run the frozen set through the judge, build the comparison table, and read the disagreements before you trust a single score.
Knowledge check
Final check
Finish the article by checking the ideas you just learned.
References
Want a more structured LLMOps path?
Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.
Large Language Models Starter Pack
A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.
- 227-page Illustrated PDF edition
- 12 guided LLM engineering chapters
- Visual concept diagrams
- Self-assessment quizzes
- Bonus deep-dive sections
- Prompt design, structured output, context windows & RAG pipelines
- Agents, tool calling, prompt injection, evaluation & application lifecycles
Coming soon


