Skip to content
intermediate

Practice Verifying LLM-Generated Code with Tests

The function looks clean. It reads well. It handles the obvious case. Then it returns the wrong number for "1h30m", and you find out three days later.

Published 2026-10-03Updated 2026-10-0410 min read
Close-up of sandy beach texture with visible footprints, creating an organic and natural pattern.
Close-up of sandy beach texture with visible footprints, creating an organic and natural pattern. Photo by James Collington on Pexels.

The function looks clean. It reads well. It handles the obvious case. Then it returns the wrong number for "1h30m", and you find out three days later.

That gap — between code that reads correctly and code that behaves correctly — is the whole reason this drill exists. You already know the specify-run-review-test loop. Now you need reps inside it, on a task small enough that the tests, not the model's confidence, decide when you're done.

Why Plausible Code Passes Your Eyes and Fails Your Users

Reading code and running code answer different questions. Reading asks, "Does this look like it does the right thing?" Running asks, "Does it actually do the right thing?" A model trained on millions of code examples is very good at producing code that survives the first question. It is not reliably good at the second.

The dangerous failure is rarely a crash. A crash is loud, and loud failures get fixed. The dangerous failure is a function that runs, returns a value, passes a quick eyeball check, and quietly mishandles one input you never thought to try. Nothing in your editor turns red. The bug ships.

There's a second trap hiding underneath. If you ask the assistant to write the code and the tests in one pass, you can get two artifacts that agree on the same mistake. The tests inherit the implementation's assumptions because the same model wrote both. A suite like that looks like verification. It's actually a mirror.

Common mistake: Treating a green test suite as proof of correctness when the tests were generated alongside the code. Agreement between two generated artifacts is not evidence. It's a shared blind spot wearing a lab coat.

So the referee has to be grounded outside the implementation. You write the expectations. You run the code against them. The success criterion for this drill is narrow and honest: a failing test that turns green for the right reason.

Set Up the Drill: One Function, One Contract

Pick a task with a clear contract and at least one non-obvious edge case. We'll use a duration parser: a function that takes a string like "1h30m" and returns the total number of seconds.

Environment assumptions, kept deliberately boring:

  • One file for the implementation, one for the tests.
  • A standard test runner for your language. No external services, no credentials, no network.
  • No dependencies beyond the standard library.

Before you touch the assistant, write the acceptance criteria in plain language:

  1. "1h30m" returns 5400.
  2. "45s" returns 45.
  3. "2h" returns 7200.
  4. "10m30s" returns 630.
  5. An empty string returns 0.
  6. A malformed string like "abc" raises an error rather than returning a number.

That's the contract. "Done" means every one of those cases passes — including the edge cases you named, not just the happy path.

Write the Expectations Before You Ask for Code

Order of operations is the entire lesson here. Expectations written first become the referee. Expectations written after the code tend to mirror the code's internal logic, because by then you've read the implementation and your brain has quietly adopted its assumptions.

Turn each criterion into a concrete input and expected output. In Python with pytest, that looks like this:

import pytest
from duration import parse_duration

def test_hours_and_minutes():
    assert parse_duration("1h30m") == 5400

def test_seconds_only():
    assert parse_duration("45s") == 45

def test_hours_only():
    assert parse_duration("2h") == 7200

def test_minutes_and_seconds():
    assert parse_duration("10m30s") == 630

def test_empty_string():
    assert parse_duration("") == 0

def test_malformed_input_raises():
    with pytest.raises(ValueError):
        parse_duration("abc")

Notice what these assertions are derived from: the specification, not the implementation. You haven't written the implementation yet. That's the point.

Tip: Write the test names so they encode the requirement. test_malformed_input_raises tells you what behavior matters. test_case_3 tells you nothing when it fails at 2 a.m.

Knowledge check

Check your understanding

Answer this question before you continue.

You have not written the duration parser yet. Which approach keeps its tests grounded in the contract rather than in implementation assumptions?
Single Choice

Focus: Translate a stated behavior contract into tests before implementation so the tests independently judge the generated code.

Generate the Code, Then Read It Like a Skeptic

Now ask your assistant for the implementation. When it comes back, resist the urge to run it immediately. Read it first, and read it looking for the things tests won't catch.

Three passes:

Check the task boundary. Does the code parse duration strings, or did it build something adjacent — a general time parser that also handles days and weeks? Scope creep in generated code is common and it hides bugs in the parts you didn't ask for.

Scan for silent assumptions. Default values, swallowed exceptions, off-by-one handling, unstated input formats. Does it accept "1H30M" in uppercase? Does it strip whitespace? Does it treat "90m" and "1h30m" the same way? Some of these are fine. The problem is when they're unstated.

Write down your prediction. Before you run anything, name which test you expect to fail and why. A prediction you can be wrong about is a prediction that teaches you something. A prediction you never make teaches you nothing.

Keep review findings separate from verification findings. Naming, style, and security are review concerns. Whether the function returns 5400 for "1h30m" is a verification concern. Both matter, but only one of them is what the tests decide.

Knowledge check

Check your understanding

Answer this question before you continue.

An assistant returns a duration parser that also accepts days, though the requested contract covers only hours, minutes, and seconds. What is the most useful next step before running the tests?
Scenario Interpretation

Focus: Review generated code for scope and unstated assumptions, then make a testable prediction before executing it.

Run the Tests and Read the Failure Honestly

A fixed contract feeds independently written tests, which check generated code. A failing test leads to diagnosis and a code change, then the full suite runs again; the loop ends at a passing suite, while the contract remains unchanged.
Keep expectations independent of the implementation; diagnose failures and re-run the full suite after each correction.

Here's a flawed implementation that produces a controlled, reproducible failure. Save it as duration.py:

import re

def parse_duration(text):
    if not text:
        return 0
    total = 0
    for value, unit in re.findall(r"(\d+)([hms])", text):
        value = int(value)
        if unit == "h":
            total += value * 3600
        elif unit == "m":
            total += value * 60
        else:
            total += value
    return total

Run the suite with pytest -v. You'll get a failure — that's the design of this drill.

Capture the exact failure: which test, which input, expected versus actual. Then classify it. There are only three possibilities:

ClassificationWhat it meansWhat to do
Wrong specificationYour expectation was incorrect or underspecifiedFix the test, re-run
Wrong implementationThe code doesn't match the contractFix the code
Wrong testThe assertion doesn't test what you meantRewrite the assertion

The failure here is test_malformed_input_raises. The implementation returns 0 for "abc" instead of raising. That's a wrong-implementation failure, and it's exactly the kind of thing that passes an eyeball check. The code "handles" bad input by ignoring it. Silent, plausible, wrong.

Check whether the failure landed where you predicted. If it did, your mental model of the code is holding up. If it didn't, your model is still incomplete, and that's worth knowing before you trust anything else about the function.

Warning: Resist the reflex to paste the error back into the assistant and accept whatever comes next. Diagnose before you delegate. If you don't know why the test failed, you can't tell whether the fix addressed the cause or just the symptom.

Knowledge check

Check your understanding

Answer this question before you continue.

With the flawed implementation shown in this section, what does `parse_duration("abc")` return?
Output Prediction

Focus: Predict the behavior of the flawed parser on malformed input by tracing what its regular-expression loop processes.

Verify the Correction Instead of Trusting It

You've diagnosed the failure. Now ask for a fix — or write it yourself, which is often faster for a bug this small. A general correction adds a validation step before parsing:

import re

def parse_duration(text):
    if not text:
        return 0
    if not re.fullmatch(r"(\d+[hms])+", text):
        raise ValueError(f"Invalid duration: {text}")
    total = 0
    for value, unit in re.findall(r"(\d+)([hms])", text):
        value = int(value)
        if unit == "h":
            total += value * 3600
        elif unit == "m":
            total += value * 60
        else:
            total += value
    return total

Re-run pytest -v. All six tests should pass now.

When the corrected code comes back, three checks:

Re-run the full suite, not just the failing test. A fix that makes one test pass while breaking another is not a fix. It's a trade you didn't agree to.

Confirm the fix addresses the cause, not the symptom. If the new code contains a branch that special-cases "abc" specifically, that's a red flag. The requirement was "malformed input raises an error," not "the string abc raises an error." A hardcoded branch that satisfies the test input without generalizing is the most common way generated fixes cheat.

Add one new case the fix should also satisfy. If the correction generalized, "xyz" should also raise. If it didn't, you'll find out now instead of in production.

One more thing to watch: if the assistant changed the test instead of the code, treat that as a specification change. Sometimes it's correct — maybe your expectation was wrong. But it needs explicit review, because "the test now passes" and "the behavior is now correct" are not the same sentence.

Knowledge check

Check your understanding

Answer this question before you continue.

A correction makes the `"abc"` test pass by adding a special case for that exact string. Which verification step best checks whether it fixed the stated requirement rather than only the example?
Debugging

Focus: Verify that a correction handles malformed inputs generally and preserves the rest of the contract.

Break Your Own Referee

A passing suite is evidence, not proof. Here's a fast way to find out whether your tests can actually detect faults.

Plant a deliberate bug in the working code. Change * 3600 to * 360 in the hours calculation, or make the empty-string case return 1. Run the suite. If nothing fails, your tests are decoration.

This matters more with generated code than with hand-written code, because of the self-confirmation loop: code and tests from the same source can share the same blind spot. The model that forgot to handle empty input is the same model that might forget to test it.

If you want to go further, two directions are worth knowing about:

  • Property-based testing defines invariants against the specification rather than the implementation — for example, "parsing any valid duration string and re-serializing it should round-trip." Tools like Hypothesis (Python) or fast-check (JavaScript) generate inputs you wouldn't think to write.
  • Mutation testing injects faults automatically and counts how many your suite catches. Tools like mutmut or PIT report a kill score: the fraction of planted bugs your tests detected.

You don't need either for this drill. But know the boundary: tests cover behavior someone wrote down. They cannot verify every production behavior that depends on the change. A green suite means "the things I thought to check are correct." It does not mean "this code is correct."

Common Mistakes in This Loop

Four failure modes show up again and again. Learn to recognize them in yourself.

Accepting a green suite that was generated alongside the code. If you didn't write the expectations independently, you verified nothing. You watched two artifacts agree.

Re-prompting on every failure instead of reading it first. The failure message contains the diagnosis. Skipping it means you're guessing at fixes and hoping one sticks.

Letting the assistant rewrite the test to match the implementation. This inverts the entire loop. The test is the referee. When the referee starts taking instructions from the player, the game is over.

Treating a passing test as proof of correctness. It's evidence about one specified behavior. That's valuable. It's not the same as proof.

The Rule You Keep

Expectations first. Run before you trust. A fix is only verified when the failing case passes for the right reason and the rest of the suite still holds.

That's the loop, and it doesn't change when the code gets bigger. Your next rep: take a real function from your own project — something small, something with a contract you can state in a few sentences. Write three acceptance cases before you ask the assistant for anything. Include one edge case the happy path would miss. Then run the same sequence: generate, read skeptically, run, diagnose, verify.

The bottleneck in AI-assisted development is not generation speed. Models will keep getting faster at producing plausible code. The constraint that actually determines whether you ship something that works is your capacity to verify what came back. Every rep you take at this loop compounds. Build the referee first, and the code generation becomes what it should have been all along: a fast draft that has to earn its way past your tests.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

An assistant generates both a function and tests, and every test passes. Which conclusion is best supported by the article?
Question 1 of 2Misconception Check

Focus: Recognize why tests written independently of generated implementation are needed to avoid shared blind spots.

Your duration-parser suite is green. You change the hours multiplier from `3600` to `360` and rerun it. What does this check tell you if the suite still passes?
Question 2 of 2Scenario Interpretation

Focus: Use a deliberately planted fault to assess whether the test suite detects behavior the contract requires.

References

  1. Demystifying evals for AI agents \ Anthropicwww.anthropic.com
  2. Reviewing AI-Generated Code: A Verification Discipline for the Loop | Augment Codewww.augmentcode.com
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.