Practice Checking LLM Claims Against Audio and Transcript Evidence
The transcript reads clean. The summary reads confident. The wrong name sails through both, and nothing in the text ever flinches.

Key topics
The transcript reads clean. The summary reads confident. The wrong name sails through both, and nothing in the text ever flinches.
That is the failure this drill is built to catch. You are going to audit a handful of claims about a short recording, judge each one against the actual audio, and learn to tell the difference between "the transcript says so" and "the recording proves it."
Why a Fluent Transcript Is Not Evidence
An audio summary is not one step. It is a chain: audio → transcript → summary. Each stage can only work with what the previous stage handed it.
Here is the part that trips people up. When a speech-to-text model mishears a name, it does not flag the mistake. It produces a clean, grammatical sentence with the wrong name in it. The summarizer downstream receives that sentence as fact, compresses it faithfully, and hands you a summary that is confident, readable, and wrong. The error entered early and traveled silently.
This is why checking a summary against its transcript is not verification. You are comparing one link in the chain to the next link. To actually verify, you have to reach back to the source: the audio itself.
Names, numbers, units, and speaker attribution are the highest-risk items. A small acoustic slip changes the meaning completely — "fifteen" becomes "fifty," a commitment gets assigned to the wrong person, a dollar figure loses its unit.
If you have read about how multimodal models handle different input types, you already know the core distinction: capability is not reliability. A model that can transcribe audio is not a model that transcribes audio correctly every time. This exercise is where that distinction becomes a habit instead of a fact you nod at.
Knowledge check
Check your understanding
Answer this question before you continue.
What You Need Before You Start
This drill runs on the Python 3 standard library alone. No installs, no API keys, no network calls.
You need the fixture set:
- A short WAV file you can play directly
- A transcript of that audio, supplied separately
- A generated summary of the transcript
claims.json— a list of claims the summary makesaudit_audio_claims.py— the script that checks your judgments
The transcript is supplied rather than generated live on purpose. This exercise is about judging evidence, not about running a model. You are the auditor here.
Note: Playback matters. You need working speakers or headphones, because the audio is the ground truth you are auditing against. If you cannot hear the recording, you cannot do this drill honestly.
When you finish, the script prints a per-claim comparison of your verdicts against the documented judgments, plus a short feedback line, and exits cleanly.
Read the Claims Before You Listen
Open claims.json and read each claim as a standalone assertion about what was said.
For each one, write down what evidence would have to exist in the audio for the claim to be true. A specific name? A number with a unit? A particular speaker making a commitment?
Then predict which claims are most likely to be fragile, and say why. Numbers and proper nouns are usually the answer, because they are the items where a small mishearing produces a large meaning change.
This is the same instinct as reading a test before you run code. You learn more from the output when you already know what you expected. Forming a prediction first gives the audio something to contradict.
Listen, Then Mark Each Claim
Now play the WAV. For each claim, assign one of three verdicts:
| Verdict | Meaning |
|---|---|
| Supported | The audio and transcript agree, and the claim matches them |
| Contradicted | The evidence says something different |
| Uncertain | You cannot tell |
Resist inventing a fourth verdict. Three is enough.
Cite the evidence for every verdict — the audio moment, the transcript line, or both. A verdict without a citation is a guess wearing a label.
Uncertain is a real answer, not a failure. Clipping, background noise, crosstalk, or a claim the recording simply never addresses all justify it. If you cannot hear it, say so.
Watch for the trap that makes this whole exercise worth doing: a claim the transcript supports but the audio does not. That is exactly the case a transcript-only check misses, and it is the one you are here to find.
Knowledge check
Check your understanding
Answer this question before you continue.
Write Your Answers in the Right Shape
The script needs a specific file format. Here is the minimal structure for answers.json:
{
"answers": [
{
"claim_id": "c1",
"verdict": "supported",
"evidence": "audio 0:04 and transcript line 3",
"action": "verify"
},
{
"claim_id": "c2",
"verdict": "uncertain",
"evidence": "audio 0:11 is clipped; transcript line 7 is ambiguous",
"action": "clarify"
}
]
}
Each entry needs four fields: the claim's ID from claims.json, your verdict, a short citation of the evidence you used, and your chosen action. The claim_id values must match the IDs in claims.json exactly — the script pairs them by ID, and a typo will show up as a missing claim rather than a wrong verdict.
If the fixture includes a template file, copy it and fill in the blanks rather than typing the structure from scratch. That removes the most common source of beginner errors.
Knowledge check
Check your understanding
Answer this question before you continue.
Decide: Verify or Clarify
Judging a claim is half the job. The other half is choosing an action.
Verify when the claim is consequential and the evidence is checkable. Go back to the audio, re-listen to the segment, or ask the speaker.
Clarify when the input itself is the problem. The audio is unusable, the speaker is ambiguous, or the claim depends on information that was never in the recording.
Do not verify everything. Rank by consequence. A wrong meeting date matters more than a wrong filler word. My rule is simple: if the claim would change a decision, it needs audio-level evidence, not transcript-level evidence.
Knowledge check
Check your understanding
Answer this question before you continue.
Run the Audit Script
From the fixture directory, run:
python audit_audio_claims.py claims.json answers.json
The script compares your verdicts against the documented judgments and reports where you diverged.
Read the divergences as evidence about your reading, not as a grade. The documented judgments are a fixture-specific key, not an oracle. They tell you where the intended discrepancy lives; they do not overrule what you actually heard. If the audio genuinely does not support a decisive answer, your "uncertain" is the more honest verdict, and the disagreement is worth investigating rather than erasing.
The documented consequential discrepancy is the one to find — a claim the summary states plainly that the audio does not support. That is the case the fixture was built around, and finding it is the point of the drill.
If the script errors, check the JSON shape and file paths first. A malformed answers.json is the most common cause.
Change One Thing and Run It Again
Now break it on purpose.
Edit one transcript line to introduce a plausible error. Swap a name. Change a number. Drop a unit. Re-run the audit.
Watch which verdicts flip. The claim text did not change. The evidence did. That is the entire lesson in one run: verdicts are a function of the evidence, not of how confident the claim sounds.
Then edit one claim instead. Notice how a small rewording can move a claim from supported to uncertain. Keep a short note of what changed and what flipped. That note is the beginning of a reusable audit habit for your own recordings.
When This Method Breaks Down
A transcript is a model's output, not a recording of truth. Treating it as infallible defeats the exercise.
If playback fails, the audio is clipped, or the recording is too noisy, the correct verdict is uncertain — not a guess dressed up as an answer.
Absent evidence is not evidence of absence. A claim the recording never addresses is uncertain, not contradicted. Those are different things, and confusing them will make you wrong in both directions.
This drill checks a handful of claims in a short clip. It does not scale to hour-long recordings without a different workflow. The habit that does scale is the question underneath all of it: which stage produced the text you are reading, and can you reach past it?
Where to Take This Next
The summary is the last link in a chain. The only way to audit it is to walk back up the chain to the audio.
Run this drill on a recording of your own — a meeting, a voice memo, an interview. Mark uncertainty honestly. Cite evidence. Verify only what would change a decision.
Then extend the fixture: add a second claim that depends on a number, and see whether your ear catches what the transcript smoothed over.
Knowledge check
Final check
Finish the article by checking the ideas you just learned.
References
Want a more structured LLMOps path?
Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.
Large Language Models Starter Pack
A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.
- 227-page Illustrated PDF edition
- 12 guided LLM engineering chapters
- Visual concept diagrams
- Self-assessment quizzes
- Bonus deep-dive sections
- Prompt design, structured output, context windows & RAG pipelines
- Agents, tool calling, prompt injection, evaluation & application lifecycles
Coming soon


