
Build a Small LLM Evaluation Harness
You changed one line in the prompt, ran three examples, and they looked fine. So you shipped. Two weeks later a customer finds the case you never re-ran —…
Read tutorialMethods for measuring, testing, comparing, and interpreting LLM or LLM-application behavior against explicit criteria.
Tagged articles
38 articles in this tag.

You changed one line in the prompt, ran three examples, and they looked fine. So you shipped. Two weeks later a customer finds the case you never re-ran —…
Read tutorial
A clean pipeline that answers confidently and wrongly is not a mystery. It is a trace you have not read yet.
Read tutorial
Six configurations sit on the whiteboard. Someone reaches for a single score to rank them, and the interesting part of the decision disappears.
Read tutorial
That is the real beginner moment, and it does not get solved by reading one more "which AI is best" listicle. Those lists go stale the moment a new model…
Read tutorial
Three tabs open. A leaderboard screenshot you saved last week. A nagging feeling that you still don't know which one to actually use.
Read tutorial
Your validation step passes. Retrieval returns documents. The model answers. Verification approves. Four green checkmarks, and the feature still fails on…
Read tutorial
The reranker "feels" better. The top result looks more relevant than it did before. And yet you cannot say whether the stage earned its latency, because…
Read tutorial
Your demo works. You type in a question, the app answers, and it looks impressive. Then you try a slightly different question, and the answer falls apart.…
Read tutorial
Your RAG system just produced a confident, polished, and completely wrong answer. The natural instinct is to blame the model—swap the prompt, change the…
Read tutorial
Frameworks change. The model underneath them doesn't. Learn the mechanism first, and every new tool becomes just another wrapper around something you…
Read tutorial
You run a few prompts. The answers are sharp, detailed, exactly what you wanted. Three for three. The system is ready, you conclude. Time to ship it.
Read tutorial
An LLM judge will grade your outputs instantly, cheaply, and at scale. It will also quietly reward verbose answers, favor responses that sound like itself,…
Read tutorial
A confidence score is a promise about a population. Most teams read it as a promise about the answer in front of them.
Read tutorial
You tweak a prompt, run a few examples, and the outputs look better. You're fairly sure the change helped—but you can't prove it. Run the same examples…
Read tutorial
Two teams evaluate the same model on the same task and report 82% and 71%. Neither team is lying. Neither team is incompetent. They are running two…
Read tutorial
Two teams run the same evaluation suite on the same two systems. One publishes a table where System A wins. The other publishes a table where System B…
Read tutorial
Version B scores 74%. Version A scored 71%. Same 40 cases, same rubric, and someone in the thread types the sentence that ends the discussion: "So B is…
Read tutorial
A threshold is not a quality setting you install. It is a bet you place on a population.
Read tutorial
A judge that agrees with humans 90% of the time can still be useless. Here is how to tell the difference.
Read tutorial
Here's a scenario I see constantly: an application sends every request to one frontier model. Summarize this email in two sentences? Frontier model.…
Read tutorial
You tweak a prompt, one case improves dramatically, you ship it, and three unrelated cases silently degrade. That is the classic trap of LLM development: a…
Read tutorial
A suite that passes every case can still fail in production. Not because the model is weak, but because the cases never mapped to the task space you…
Read tutorial
Your eval score went up. The change shipped. Nobody asked whether the test set was ever clean.
Read tutorial
Your judge reports a 92% pass rate. The dashboard is green. Then a reviewer reads ten of those passes and disagrees with four of them.
Read tutorial
Most teams add hybrid retrieval because a vendor page promised better accuracy, then never check whether it helped their own corpus. This exercise forces…
Read tutorial
The cheapest route on the pricing page is often the most expensive one in production.
Read tutorial
You swap a model to a 4-bit build, run the same prompt, and get a slightly different answer. Now what? You cannot tell whether quantization degraded the…
Read tutorial
You run the same prompt five times and get five answers that mean roughly the same thing but never match word for word. Now what? You cannot tell whether…
Read tutorial
Your retriever finally returns the right documents. The chunks look relevant. The context window is full of plausible evidence. And the model still hands…
Read tutorial
A citation marker is a promise. The sentence says this passage backs me up. Most RAG evaluations never check whether the promise holds — they check whether…
Read tutorial
Two retrievers return the same relevant chunk. One puts it at rank 1; the other buries it at rank 4. A single "accuracy" number calls them identical. The…
Read tutorial
Your fallback chain looks correct in review. Then the primary stalls at 900ms, the backup is rate-limited, and nobody can say whether the request…
Read tutorial
A chat assistant repeats a detail from ten turns ago, then forgets the order number you gave it two turns ago. Nothing about the model changed between…
Read tutorial
Retrieval misses rarely announce their cause. The answer was in the document, the query was reasonable, and the model still returned something useless. The…
Read tutorial
You add a rewrite step. The demo answers the question. You feel clever. Then someone asks a slightly different question and the whole thing falls apart,…
Read tutorial
A reusable prompt template is an interface: fixed instructions plus a variable slot whose values you do not control. One clean input exercises the…
Read tutorial
A filter can be correct and still delete your best evidence. Here is how to watch it happen in twenty lines of NumPy.
Read tutorial
You added three examples to your prompt. The output looked better. You shipped it.
Read tutorial