The first three articles in this series were about the grand stuff — trillion-token corpora, scaling economics, synthetic generation. This one is about the part I’ll actually do with my own hands: building a dataset to fine-tune a model for a specific purpose. It’s the least published, most artisanal corner of the field — the knowledge lives in blog posts, Discord servers, and hard-won failure — so consider this my field guide, assembled from the best of what I’ve found and tried to systematize.
First decision: what fine-tuning is actually for
The most expensive mistakes happen before any data is collected, by fine-tuning for the wrong reason. The clarifying frame (which LIMA’s superficial-alignment result backs theoretically): fine-tuning is superb at teaching behavior — format, style, tone, task procedure, domain vocabulary, when to refuse, how to structure output — and mediocre-to-poor at injecting knowledge, especially fresh or frequently-changing facts. The model already knows what it knows from pretraining; SFT mostly selects and shapes how that knowledge is presented.
So the standing rule of thumb: facts that change → retrieval (RAG); behavior that repeats → fine-tuning. They compose beautifully — fine-tune the model to use retrieved context in your exact format — but substituting one for the other fails predictably: fine-tuning on your product docs won’t make the model reliably quote current prices, and RAG alone won’t make it consistently produce your report format. If the goal doesn’t survive being phrased as “I want the model to behave differently on inputs like these,” reconsider the whole project.
The anatomy of an SFT example — and where quality actually lives
An SFT dataset is conversations: (system?, user, assistant) turns, usually in a chat template, with loss computed only on the assistant’s tokens. Simple container; the craft is entirely in the contents, and it concentrates in four places:
Inputs must look like production. The single most common failure I’ve read about: training on clean, well-phrased prompts, deploying against messy real users — typos, missing context, half-Thai half-English, pasted logs. The training inputs should be drawn from (or faithfully imitate) the ugly true distribution, including the hard and weird cases, not the ones that are pleasant to write.
Outputs must be exactly what you want — every single one. LIMA’s lesson, operationalized: the model learns the average of your demonstrations. One sloppy example doesn’t just fail to help; it teaches sloppiness as an acceptable mode. Every response in the set should pass the test “would I be happy if the deployed model produced precisely this?” — same format, same tone, same length discipline. Style consistency across examples matters more than the elegance of any one of them.
Coverage beats volume. A thousand diverse, curated examples routinely outperform fifty thousand scraped ones for behavioral tuning. The diversity axes to deliberately enumerate: task variants, input formats, difficulty levels, edge cases, and the negative space — inputs the model should refuse, questions it can’t answer from the given context (train it to say so — this is where a chunk of hallucination behavior is won or lost), adversarial phrasings.
Provenance is not paperwork. Track where every example came from and its license. Datasets generated from other models’ outputs may carry terms-of-use restrictions; scraped answers carry copyright; user data carries privacy law. The time to know is before training, not after shipping.
Where do examples come from? In practice, a funnel that combines everything from the last article: real logged interactions (best distribution, needs cleaning and consent), human-written seeds for the core behaviors (small, load-bearing), and synthetic expansion — a strong model generating variations over your taxonomy of cases — followed by filtering, because generation is cheap and judgment is the product. A pattern that consistently works: humans write the rubric and a few dozen gold examples; a strong LLM drafts hundreds more; a verifier (rules, tests, or LLM-judge with the rubric) filters; a human spot-checks the survivors. Most modern datasets, including famous open ones, are exactly this kind of human-machine sandwich.
Preference data: teaching “better,” not just “correct”
SFT shows the model one good answer. Preference data — pairs of (chosen, rejected) responses to the same prompt — teaches a direction: away from this, toward that. With DPO-family training (my RLHF article covers the mechanics), this is how you fix the residual behaviors SFT can’t quite pin down: verbosity, hedging, format drift, sycophancy, subtle wrongness that only shows in contrast.
The craft points that recur across every writeup I trust: the rejected response should be plausible — an answer the current model actually produces — not a strawman, or the training signal is trivially easy and changes nothing. The best source of rejected examples is therefore your own model’s current failures on real prompts, with the chosen side being a corrected version. Watch for label artifacts: if “chosen” is systematically longer, you’re training a length preference (a real, documented pathology of preference pipelines — models balloon in verbosity because labelers unconsciously reward length). And small, clean preference sets beat large noisy ones even more sharply than in SFT, because every pair is a high-leverage gradient about taste.
Evaluation data: the part everyone builds last and should build first
Here’s the practice that separates disciplined projects from vibes-driven ones: hold out an eval set before training anything, drawn from the same real distribution, and define how you’ll score it. Automatic checks where possible (exact match, format validators, unit tests for code); an LLM-judge with a written rubric for the fuzzy dimensions; a small human-reviewed slice to calibrate the judge. Two non-negotiables the whole field keeps relearning: decontaminate — make sure eval prompts (or near-paraphrases) aren’t sitting in the training set, which silently happens constantly, especially with synthetic expansion that may regurgitate seeds; and evaluate the behaviors you didn’t train — a fine-tune can sharpen your task while quietly degrading general instruction-following or safety behavior (catastrophic forgetting is mostly avoidable with mixed-in general data, but only if you’re measuring it).
The loop that actually improves models, then, isn’t train-once: it’s error-driven iteration. Train → run the eval → read the failures (actually read them — categorize by hand) → author or generate targeted examples for exactly those failure modes → retrain. Every serious practitioner writeup converges on the same claim: ten examples aimed at a diagnosed failure beat a thousand generic ones. The dataset becomes less like a pile and more like a changelog of fixed bugs — which is, I think, the correct mental model for the whole artifact.
The meta-lesson of the whole series
Four articles ago I started with pipelines that demolish the raw web; then the quality-vs-quantity economics; then machines writing their own curriculum; now the handcraft of a thousand curated examples. The through-line is one idea at every scale: training data is not a found resource — it is a designed artifact, and the design decisions (what counts as good, what gets filtered, what gets repeated, what gets refused) are quietly the most consequential decisions in machine learning. Models are mirrors with a loss function. Whatever care, judgment, and taste you encode in the data comes back out as capability — and whatever you didn’t, comes back out too.
Next series: LoRA — how to actually train on all this data without renting a datacenter.