Skip to content
beginner

Practice Reviewing a Small Fine-Tuning Dataset

You generated forty examples. The file is open. Every row looks fine, and that is the problem — "looks fine" is not a review. A dataset is not a pile of…

Published 2026-10-03Updated 2026-10-0412 min read
Stunning sunset with a crescent moon, rich orange and red hues create a dramatic sky view.
Stunning sunset with a crescent moon, rich orange and red hues create a dramatic sky view. Photo by Oleksiy Yeshtokyn,🌻🇺🇦🌻 on Pexels.

You generated forty examples. The file is open. Every row looks fine, and that is the problem — "looks fine" is not a review. A dataset is not a pile of examples; it is a written specification of how you want the model to behave. Most beginners never read their own specification. They ship it to a training run and then act surprised when the model behaves like a committee that never met.

So let's do the cheap version first. We will review one small synthetic dataset against an explicit behavior objective, hunt for three classes of defects, and fix a bounded subset. No GPU, no API key, no training run.

What You Need Before You Start

This exercise assumes you already understand that fine-tuning adjusts behavior rather than adding knowledge, and that prompting or retrieval is often the better first move. That decision is behind us. If it isn't, go settle it before you spend an afternoon on data review.

Here is the full kit:

  • A text editor.
  • A way to read JSONL or a simple CSV — a spreadsheet works fine.
  • Thirty to sixty minutes.
  • One sentence you write yourself.

That last item is the real prerequisite. Before you touch a single row, write a behavior objective: one or two sentences stating what the model should do, for whom, in what format, and what it should refuse.

Why does the objective come first? Because without it, "inconsistent" and "missing coverage" have no definition. Inconsistency is not a property of a row. It is a row contradicting your stated intent. The objective is the ruler. No ruler, no measurement — just vibes.

Common mistake: Reviewing examples against your taste instead of against a written objective. Taste drifts between rows. A written objective doesn't.

The Practice Dataset: A Support-Reply Task

Here is our objective:

Behavior objective: A billing-support assistant that answers customer billing questions in a calm, professional tone, always replies in this shape — a one-line acknowledgment, then a direct answer, then a next step — and escalates any refund request over $100 to a human agent instead of answering it. It must not state specific policy details (retry windows, fee rules, timelines) that are not present in the customer's message.

The format is an instruction/input/output triple. The same review method applies to chat-style message lists; you just read the message array instead of three fields.

{"instruction": "Answer the customer's billing question.", "input": "Why was I charged twice this month?", "output": "Thanks for flagging this. A duplicate charge usually means a retry after a failed payment. I've noted your account for review, and you'll see a correction within 3-5 business days."}
{"instruction": "Answer the customer's billing question.", "input": "Can I get a refund for the $240 annual plan?", "output": "Thanks for reaching out. I've escalated your refund request to a human agent who will follow up within one business day."}
{"instruction": "Answer the customer's billing question.", "input": "I want a refund for the $40 add-on.", "output": "Thanks for reaching out. I've escalated your refund request to a human agent who will follow up within one business day."}
{"instruction": "Answer the customer's billing question.", "input": "How do I update my card?", "output": "Thanks for flagging this. You can update your card under Settings > Billing > Payment Methods. Once saved, the new card is used on your next invoice."}
{"instruction": "Answer the customer's billing question.", "input": "My invoice shows a $15 late fee I don't recognize.", "output": "Thanks for flagging this. Late fees are applied when a payment fails and isn't retried within 7 days. I've noted your account for review, and you'll see a correction within 3-5 business days."}
{"instruction": "Answer the customer's billing question.", "input": "Can I get a refund for the $240 annual plan?", "output": "Thanks for reaching out. Refunds over $100 are handled by our billing team. I've escalated this and someone will follow up within one business day."}
{"instruction": "Answer the customer's billing question.", "input": "This is ridiculous, I've been overcharged three times!", "output": "Thanks for flagging this. I understand the frustration. I've noted your account for review, and you'll see a correction within 3-5 business days."}
{"instruction": "Answer the customer's billing question.", "input": "Do you offer a student discount?", "output": "Thanks for flagging this. Yes, students get 20% off with a valid .edu email. You can apply it under Settings > Billing > Discounts."}
{"instruction": "Answer the customer's billing question.", "input": "Why did my plan change from monthly to annual?", "output": "Thanks for flagging this. Plan changes only happen when you confirm them, so this may be a display issue. I've noted your account for review, and you'll see a correction within 3-5 business days."}
{"instruction": "Answer the customer's billing question.", "input": "I need a refund for the $240 annual plan, I was charged after cancelling.", "output": "Thanks for reaching out. I've escalated your refund request to a human agent who will follow up within one business day."}
{"instruction": "Answer the customer's billing question.", "input": "How do I download my invoices?", "output": "Thanks for flagging this. You can download invoices under Settings > Billing > History. Each invoice has a PDF link."}
{"instruction": "Answer the customer's billing question.", "input": "Can I get a refund for the $40 add-on?", "output": "Thanks for reaching out. I've escalated your refund request to a human agent who will follow up within one business day."}
{"instruction": "Answer the customer's billing question.", "input": "Why was I charged twice this month?", "output": "Thanks for flagging this. A duplicate charge usually means a retry after a failed payment. I've noted your account for review, and you'll see a correction within 3-5 business days."}
{"instruction": "Answer the customer's billing question.", "input": "What's the difference between the Pro and Team plans?", "output": "Thanks for flagging this. Pro is for individuals; Team adds shared billing and seats. You can compare them under Settings > Billing > Plans."}
{"instruction": "Answer the customer's billing question.", "input": "I was charged $240 for the annual plan but I only signed up for monthly.", "output": "Thanks for flagging this. Annual plans are billed once per year, so the $240 charge is expected. I've noted your account for review, and you'll see a correction within 3-5 business days."}

Fifteen rows. Read them all — that is the point of keeping the set small. Synthetic examples carry a specific risk profile: they tend to be too uniform, too polite, and too confident. You can already feel it in the repeated "Thanks for flagging this."

Defect One: Inconsistent Examples

Inconsistency comes in three flavors:

  • Format drift — different output shapes for the same task.
  • Tone drift — some replies formal, some casual.
  • Policy drift — the same request answered two different ways.

Look at rows 2 and 6. Both are refund requests over $100. Row 2 escalates. Row 6 also escalates, but with different wording — that's mild format drift, not fatal. Now look at row 15: a $240 charge, no escalation, and the assistant asserts the charge is "expected." That is policy drift. The same class of request gets two different behaviors.

Here is the mechanism. The model learns the distribution it is shown. If some refund replies escalate and others answer directly, the model learns to be unpredictable on refunds. It does not average the two behaviors into a sensible middle. It samples from both.

A practical detection pass: sort the examples by input type and read the outputs as a column. Contradictions become visible when similar inputs sit next to each other. Row-by-row reading hides them; column reading exposes them.

Note: Some variation is legitimate. Different inputs should produce different outputs. The defect is variation your objective does not authorize.

Knowledge check

Check your understanding

Answer this question before you continue.

The review finds row 15 saying a $240 annual-plan charge is expected, while the objective calls for escalation of refund requests over $100. What revision does the article make to address this policy drift?
Debugging

Focus: Identify how the article recommends correcting a row that contradicts the stated behavior objective.

Defect Two: Leakage and Unanswerable Examples

Leakage means the output contains a fact, number, name, or policy detail that does not appear in the input or in the system's allowed knowledge. The model is being trained to invent that detail confidently.

Row 5 is the clean example. The customer asks about a $15 late fee. The output states late fees apply "when a payment fails and isn't retried within 7 days." Where did the 7 days come from? Not the input. Not the objective. The example is teaching the model to hallucinate a specific policy number with total confidence.

The related defect: examples where the input is genuinely ambiguous but the output answers as if it were clear. Row 9 asks why a plan changed. The output says it "may be a display issue" — a guess dressed as an answer. That teaches overconfidence on underspecified requests.

Detection pass: cover the output. Ask whether the input alone contains enough information to produce it. If not, the row is either leaking or under-specified.

Warning: Leakage is easy to spot in fifteen rows and brutal at five thousand. That is exactly why you build the habit small first.

Knowledge check

Check your understanding

Answer this question before you continue.

A customer asks about an unfamiliar $15 late fee, and the reply claims late fees follow a seven-day retry rule. According to the article's revision approach, what should the reviewer do?
Scenario Interpretation

Focus: Recognize and revise an output that supplies a policy detail absent from both the input and the objective.

Defect Three: Coverage Gaps

Coverage is about the distribution of inputs, not the count of rows. Fifteen rows can cover a task well or cover one narrow slice fifteen times.

Build a coverage table. List the input categories your objective implies, then tally.

Input categoryRowsCount
Duplicate / unexpected charge1, 132
Refund request (over $100)2, 6, 10, 154
Refund request (under $100)3, 122
How-to / self-service4, 112
Plan / pricing question8, 142
Angry or frustrated customer71
Ambiguous or under-specified91

Two problems jump out. First, over-representation: four of fifteen rows are the same over-$100 refund case, and two of those contradict each other. Second, the objective's format rule — acknowledgment, direct answer, next step — is never tested against an input that resists a direct answer. Every row here is answerable, so the dataset never shows the model what to do when it cannot answer.

The consequence of a gap is not silence. The model will still produce something for uncovered inputs — and that something will be whatever the base model defaults to, not what your objective specifies. A gap is not a hole. It is an unwritten rule.

Knowledge check

Check your understanding

Answer this question before you continue.

The dataset has no example showing what the assistant should do when it lacks enough information to answer directly. Which review action best addresses this coverage gap?
Comparison Reasoning

Focus: Use a coverage review to identify and address a missing case where the assistant cannot directly answer.

Revise a Bounded Subset

A flowchart starts with the behavior objective, branches into three checks labeled consistency, leakage, and coverage, then converges on logging a confirmed defect and revising at most five rows. A separate final step points to held-out evaluation.
Use the objective to guide three focused checks, and keep revisions bounded; a cleaner dataset still needs evaluation.

Set the bound before editing: revise at most five rows, and only rows you can tie to a named defect. Unbounded editing is how a review turns into a rewrite, and a rewrite has no audit trail.

For each revised row, record three things: the defect, the change, and the objective clause the change serves.

Revision 1 — policy drift (row 15).

Before: the assistant tells the customer the $240 charge is "expected." After: the assistant escalates, matching rows 2 and 6. Objective clause served: "escalates any refund request over $100."

Revision 2 — leakage (row 5).

Before: the output cites a 7-day retry window that appears nowhere in the input. After: the output acknowledges the fee and escalates for review without inventing the policy detail. Objective clause served: "must not state specific policy details that are not present in the customer's message."

Now the harder discipline. Leave row 9 alone. It looked odd — the assistant guesses "display issue" — but under the objective, an ambiguous input still deserves a next step, and the row provides one. Not every strange example is a defect. If you cannot state the defect and the objective clause it violates, do not edit the row yet. Go sharpen the objective instead.

Finally, add one new row to close the format gap — an input the assistant genuinely cannot resolve directly:

{"instruction": "Answer the customer's billing question.", "input": "I think I was charged the wrong amount but I'm not sure which invoice it was.", "output": "Thanks for reaching out. I don't have enough detail to identify the charge yet. Please share the invoice date or the last four digits of the card, and I'll look into it."}

Note what just happened: adding a row changed the distribution. That new row must be consistent with the existing format and tone, or you have created a fresh inconsistency while fixing a gap.

Knowledge check

Check your understanding

Answer this question before you continue.

Row 9's response to an ambiguous plan-change question seems odd, but the article says it provides a next step under the objective. What should the reviewer do?
Misconception Check

Focus: Apply the article's rule for deciding whether a questionable example should be edited.

What This Exercise Does Not Prove

A cleaner dataset is a better specification. It is not a better model. Training dynamics, base model choice, data volume, and evaluation design all sit between your edits and the behavior you observe.

The only way to know whether a revision helped is to hold out examples, train or re-train, and compare on the held-out set. Even then, small datasets produce noisy comparisons. Treat any improvement as a signal, not a verdict.

Three failure modes to watch for after this exercise:

  • Over-editing until the dataset is sterile and every row sounds identical.
  • Tuning to a single test case — shaping data around one example you happen to care about.
  • Treating a clean review as a finished project.

My rule: if you cannot name the defect and the objective clause it violates, do not edit the row. Sharpen the objective instead.

Your Next Move

Write your own one-sentence behavior objective. Then review ten of your own examples against it and log the defects you find — before changing anything. The log is the asset. It scales to larger datasets, it survives team handoffs, and it catches problems at the cheapest possible moment: before a training run pays for them.

The review habit is the reusable thing. The dataset is just where you practice it.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

After reviewing and revising a small dataset, what conclusion is justified by the article?
Question 1 of 2Misconception Check

Focus: Distinguish a cleaner dataset specification from evidence that training improved model behavior.

You are beginning a review of your own fine-tuning examples. Which sequence best matches the article's recommended next move?
Question 2 of 2Scenario Interpretation

Focus: Choose a first review step that records defects before making changes to examples.

References

  1. Fine-tuning best practices | OpenAI APIdevelopers.openai.com
  2. How to fine-tune: Focus on effective datasetsai.meta.com
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.

A white humanoid toy robot standing on a reflective black surface in a studio setting with a blue and pink gradient background.
beginner
8 min read

A Brief History of LLMs

Large language models look like overnight magic, but they are the visible tip of decades of compounding engineering. If you want to use, trust, and build…

Read tutorial