Build a Validated Document-Extraction Workflow
A model returns clean JSON. The pipeline writes it to a table. One wrong field quietly poisons everything downstream.

Key topics
A model returns clean JSON. The pipeline writes it to a table. One wrong field quietly poisons everything downstream.
That is the failure this tutorial exists to prevent. You already know the shape of an extraction workflow — fields, evidence, validation, review — and you already know model output is untrusted data. This is the runnable version: one small script that extracts a few fields from a sample document, checks each value against a contract and against the source text, and routes the shaky ones to a human instead of silently trusting them.
What You Are Building and What You Need
The target is a script that reads a document, extracts a small fixed set of fields, validates every field, and emits three buckets: accepted, needs-review, rejected.
Assumptions before you start:
- Documents are text-extractable. Plain text files, or a clean PDF-to-text step. If your PDFs are scans, you need OCR first, and that is a separate problem.
- You have a model that returns JSON. An API key for a hosted model, or a local model with structured output support. The example below uses a fixed response fixture so you can run the whole workflow without a live call, then swap in your model adapter.
- You keep the field set small. Three to five fields. This is not a limitation to work around — it is what keeps validation logic readable and failures diagnosable. A twelve-field schema on day one produces a review queue nobody can debug.
State your success criteria up front: every accepted row has a value, a source span, and a passing contract check. If a row cannot meet all three, it does not get accepted. That single rule is the whole design.
Define the Field Contract Before You Prompt
The contract is the source of truth. The prompt is downstream of it.
A contract is a per-field spec. For each field, write down:
| Property | Example |
|---|---|
| Name | invoice_date |
| Type | date |
| Format | ISO 8601 (YYYY-MM-DD) |
| Required | yes |
| Valid value | a real calendar date within the last 10 years |
| Normalization | parse any recognizable date format, emit ISO |
Include a normalization rule per field so validation has something deterministic to check. Currency becomes a numeric amount plus a currency code. Dates become ISO strings. Names get trimmed and whitespace-collapsed.
Then decide your missing-value convention explicitly. An absent field is a distinct outcome from a wrong field, and both differ from a field the model guessed. If you do not decide this now, you will discover it later as three different bugs wearing the same mask.
Common mistake: Writing the prompt first and reverse-engineering the schema from whatever the model happens to return. This feels faster. It also means your contract is defined by model behavior rather than by what your downstream system actually needs — and you lose the only artifact that lets you reject output without arguing about whether the model "seemed" right.
Knowledge check
Check your understanding
Answer this question before you continue.
Run the Smallest Extraction Pass
Here is the prompt shape. Give the model the contract, ask for JSON only, and require a verbatim source quote for each field alongside the value.
Extract the following fields from the document below.
Return JSON only, with this shape:
{
"invoice_date": {"value": "...", "evidence": "..."},
"total_amount": {"value": "...", "evidence": "..."},
"vendor_name": {"value": "...", "evidence": "..."}
}
Rules:
- "value" is the extracted value, normalized per the field spec.
- "evidence" is a verbatim quote from the document that supports the value.
- If a field is not present, use null for both value and evidence.
Document:
<document text here>
The evidence field is the key design choice. Asking for the supporting span turns an unverifiable answer into something you can check mechanically. Without it, you are trusting the model's confidence. With it, you have a string you can search for in the source.
Run this against one document and inspect the raw response before adding any validation. You need to see the failure modes your validator must catch. Here is a realistic one:
{
"invoice_date": {
"value": "2024-03-15",
"evidence": "Invoice Date: March 15, 2024"
},
"total_amount": {
"value": 1240.00,
"evidence": "Total Due: $1,240.00"
},
"vendor_name": {
"value": "Acme Supply Co.",
"evidence": "From: Acme Supply Co."
}
}
That looks clean. Now imagine the same response where total_amount carries the evidence "Total Due: $1,240.00" but the document actually says $1,420.00. The model produced a plausible value with a quote that does not appear in the text. That is the failure your validator exists to catch, and you will only see it if you look at raw output first.
Note: Long documents may need chunking or a pre-extraction text cleanup step. That cleanup is where a lot of silent errors enter — page numbers stripped mid-sentence, tables flattened into unreadable runs of numbers. Check your cleaned text before you blame the model.
Validate Values Against the Contract
Now add the deterministic checks. Three families:
- Presence — required field returned, not null.
- Shape — type and format match the contract.
- Plausibility — value falls in an expected range or set.
Run these before any model-based judgment. Deterministic checks are cheap, fast, and explainable. You never want to spend a model call on a value that fails a regex.
Emit a validation result per field: pass, fail with reason, or missing. The reason string is what makes the review queue usable later — "date parses but is in 1912" tells a reviewer something. "Invalid" does not.
from datetime import datetime, timedelta
def validate_field(name, spec, extracted):
if extracted is None or extracted.get("value") is None:
return {"status": "missing", "reason": f"{name} not found in document"}
value = extracted["value"]
if spec["type"] == "date":
try:
parsed = datetime.fromisoformat(value)
except (ValueError, TypeError):
return {"status": "fail", "reason": f"{name} is not a valid ISO date"}
cutoff = datetime.now() - timedelta(days=365 * 10)
if parsed < cutoff or parsed > datetime.now():
return {"status": "fail", "reason": f"{name} outside the 10-year contract window"}
if spec["type"] == "amount":
if not isinstance(value, (int, float)) or value <= 0:
return {"status": "fail", "reason": f"{name} is not a positive number"}
return {"status": "pass", "reason": ""}
The date rule now matches the contract exactly: a real calendar date within the last ten years. The earlier draft used a hard-coded parsed.year < 2000 boundary, which quietly disagreed with the stated field spec. That kind of drift is the exact bug the contract is supposed to prevent, so keep the two in sync.
Common mistake: Treating a schema-valid JSON response as a valid extraction. Well-formed output and correct output are different claims. The model can return perfect JSON containing a date that parses cleanly and is still wrong by a century.
A useful failure to demonstrate on purpose: a total that is arithmetically inconsistent with its line items. Your contract can require that the extracted total equals the sum of extracted line amounts, and that check catches errors no format rule will.
Knowledge check
Check your understanding
Answer this question before you continue.
Check Each Value Against Its Source Evidence
This is the step that separates a grounded extraction from a confident guess.
Verify the returned quote actually appears in the source text. Allow for whitespace and minor normalization differences — collapse runs of spaces, normalize quotes and dashes — but require the substance to match.
import re
def normalize(text):
return re.sub(r"\s+", " ", text).strip().lower()
def evidence_check(extracted, source_text):
evidence = extracted.get("evidence")
if not evidence:
return {"status": "fail", "reason": "no evidence provided"}
if normalize(evidence) in normalize(source_text):
return {"status": "pass", "reason": ""}
return {"status": "fail", "reason": "evidence not found in source"}
When the quote is absent, that is a strong signal the value was inferred or hallucinated. Treat it as a review case, not a silent pass.
For fields where the value is derived rather than quoted — a computed total, a normalized date — define what evidence is acceptable and check that instead. A date normalized from "March 15, 2024" should carry the original string as evidence, not the ISO form.
The mechanism is plain: the model produces text that looks like evidence. Only a comparison against the document makes it evidence.
Warning: Quote matching proves the span exists. It does not prove the span is the right answer to the field question. A model can quote a real sentence that supports a different field entirely. Evidence checking narrows the failure space; it does not close it.
Knowledge check
Check your understanding
Answer this question before you continue.
Wire the Pieces Into One Runnable Pass
The fragments above only matter if they compose. Here is the whole workflow in one script: a sample document, a fixed model response, the contract, both check layers, and the routing decision. Run it as-is, then replace MODEL_RESPONSE with a live call.
import json
from datetime import datetime, timedelta
DOCUMENT = """Acme Supply Co.
Invoice Date: March 15, 2024
Total Due: $1,240.00
Thank you for your business."""
MODEL_RESPONSE = json.dumps({
"invoice_date": {"value": "2024-03-15", "evidence": "Invoice Date: March 15, 2024"},
"total_amount": {"value": 1240.00, "evidence": "Total Due: $1,240.00"},
"vendor_name": {"value": "Acme Supply Co.", "evidence": "From: Acme Supply Co."},
})
CONTRACT = {
"invoice_date": {"type": "date", "required": True},
"total_amount": {"type": "amount", "required": True},
"vendor_name": {"type": "text", "required": True},
}
def route(record):
"""Accepted only if every field passes both checks. Otherwise review or reject."""
if record is None:
return "rejected", "unparseable model output"
field_results = {}
for name, spec in CONTRACT.items():
extracted = record.get(name)
value_result = validate_field(name, spec, extracted)
evidence_result = evidence_check(extracted or {}, DOCUMENT)
field_results[name] = {"value": value_result, "evidence": evidence_result}
# Reject only when the record is unusable: every field missing or malformed.
all_missing = all(r["value"]["status"] == "missing" for r in field_results.values())
if all_missing:
return "rejected", "no usable fields extracted"
# Review when any field fails a check or lacks evidence.
needs_review = any(
r["value"]["status"] != "pass" or r["evidence"]["status"] != "pass"
for r in field_results.values()
)
if needs_review:
reasons = [
f"{name}: {r['value']['reason'] or r['evidence']['reason']}"
for name, r in field_results.items()
if r["value"]["status"] != "pass" or r["evidence"]["status"] != "pass"
]
return "review", "; ".join(reasons)
return "accepted", ""
try:
record = json.loads(MODEL_RESPONSE)
except json.JSONDecodeError:
record = None
bucket, reason = route(record)
print(f"route: {bucket}")
print(f"reason: {reason}")
Run it. The output is:
route: review
reason: vendor_name: evidence not found in source
That is the workflow doing its job. The date and amount pass both checks. The vendor name fails evidence checking because the model invented the label "From:" — the document says Acme Supply Co. but never says From:. The value is probably correct, but the evidence is not verbatim, so the record goes to review instead of into a trusted table.
This is the boundary the earlier draft left ambiguous. The rule is now explicit and applied consistently:
- Rejected — the record is unusable. Unparseable JSON, or every field missing.
- Review — the record is interpretable but at least one field fails a check or lacks matching evidence.
- Accepted — every field passes both the contract check and the evidence check.
Change MODEL_RESPONSE to fix the vendor evidence and re-run. The route flips to accepted. Break the date to "1899-01-01" and it flips to review with a contract reason. That is the loop you will run constantly.
Knowledge check
Check your understanding
Answer this question before you continue.
Route Uncertain Cases to a Review Queue
The script above prints a route. In production, that route becomes a queue entry.
Design the review record so a human can decide fast. Include the source document, field name, extracted value, the evidence span, and the specific failed check. A reviewer should not have to open the document and re-read it to understand what went wrong.
Set the routing threshold deliberately. A strict threshold sends more work to humans; a loose one pushes risk downstream. The right setting depends on what a wrong value costs. A wrong shipping address is expensive. A wrong internal tag is annoying.
Common mistake: Routing on a single overall confidence number instead of per-field reasons. This makes review slow and hides which fields are systematically unreliable. If
vendor_namealways lands in review, that is a contract or prompt problem, not a human problem. Track per-field review rates and fix the field, not the queue.
Debug the Failures You Will Actually Hit
Keep this list nearby. These are the breakages that show up in real runs.
- Malformed or truncated JSON. Usually a prompt or max-token issue, not a model capability issue. Check your output length limit before you change models.
- Right value, wrong field. Often caused by ambiguous field names or overlapping definitions in the contract. If
invoice_dateanddue_dateboth say "date," the model will guess. - Correct-looking value with no matching evidence. The classic hallucination signature. Check whether the document text was actually passed in full — a truncated input produces confident answers about text the model never saw.
- Works on one document, fails on another. Usually a layout or text-cleanup difference, not a model difference. Diff the cleaned text between the two documents.
Debugging habit: keep the raw model response for every document. You can re-run validation without re-calling the model, which makes iterating on your checks nearly free.
One Experiment: Tighten the Contract and Watch the Review Rate Move
Change one thing at a time. Add a stricter format rule to one field, or split an ambiguous field into two clearer ones. Re-run the same sample documents and compare how many fields land in auto-accept versus review.
Interpret the result carefully:
- Review rate drops, no new failures. The contract got clearer. Good.
- Review rate drops, new wrong values appear. You loosened a check you needed. Revert.
This is the loop that matters in production. The contract and the checks are the tunable surface, not the model. Swapping models is the expensive move; sharpening a field definition is the cheap one.
Note: A handful of sample documents tells you about your checks, not about your accuracy on the full corpus. Treat the experiment as a signal about your validation logic, not a benchmark.
Where This Leaves You
The rule to carry forward: never let a model value reach a trusted table without a contract check and an evidence check, and never let an uncertain value disappear — route it.
You now own a reusable asset: the contract-plus-checks loop, running end to end on one document. The next practical step is to expand your sample set and track per-field review rates. The fields that consistently land in review are telling you where your contract is still vague. Fix those definitions first, and the review queue shrinks on its own.
Knowledge check
Final check
Finish the article by checking the ideas you just learned.
References
Want a more structured LLMOps path?
Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.
Large Language Models Starter Pack
A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.
- 227-page Illustrated PDF edition
- 12 guided LLM engineering chapters
- Visual concept diagrams
- Self-assessment quizzes
- Bonus deep-dive sections
- Prompt design, structured output, context windows & RAG pipelines
- Agents, tool calling, prompt injection, evaluation & application lifecycles
Coming soon


