Data Engineering · Part 3

Synthetic Data: Training Models on Text That Models Wrote

There's an idea that would have sounded like a joke a few years ago: take a language model, have it write its own training data, and train the next model on it.

Good Omens Studio7 min readData Engineering

There’s an idea that would have sounded like a joke a few years ago: take a language model, have it write its own training data, and train the next model on it. Surely that’s an echo chamber — photocopying a photocopy. Yet synthetic data went from taboo to arguably the most important data trend in the field: the phi models were built on it, frontier reasoning models are trained on their own generated reasoning traces, and nearly every serious fine-tuning dataset today is partly or wholly model-written. This article is my attempt to sort out when generating your own data works brilliantly, when it degenerates, and why the difference isn’t luck — it’s structure.

Why anyone reached for this in the first place

Three pressures made synthetic data inevitable:

The internet is running out. High-quality human public text is finite, and frontier training runs consume tokens faster than humanity writes them. Data, once assumed infinite, is becoming the binding constraint on the scaling curve — and when a resource becomes scarce, you start manufacturing it.

Human labeling doesn’t scale, and isn’t even best. Instruction datasets and preference labels are expensive; worse, for many tasks a strong model is already a better annotator than a bored crowdworker — more consistent, infinitely patient, pennies per example. Constitutional AI made this explicit early on: replace much of the human feedback in alignment with AI feedback guided by written principles (RLAIF), and quality holds while cost collapses.

Some data barely exists naturally. Step-by-step reasoning traces, code paired with tests, dialogues exhibiting perfect tool use, textbook explanations calibrated to a difficulty curve — the internet contains conclusions, not the thinking. If you want to train on the process, someone has to produce it — and models can.

The taxonomy: not one technique, but four

“Synthetic data” is used for genuinely different things, and conflating them is where most confusion (including mine) came from.

1. Distillation data: a stronger teacher writes for a weaker student. Prompt a frontier model to generate instructions and responses; fine-tune a small model on the output. This is the Alpaca/Vicuna lineage that ignited open-source LLMs, and — via reasoning traces — how compact models inherit chain-of-thought ability from large reasoners. It works because there’s a capability gradient: information genuinely flows downhill from a better model. The catch is a ceiling (the student approaches, rarely exceeds, the teacher on that slice) and a characteristic failure mode documented in the “false promise of imitation” critique: shallow imitation copies the teacher’s style — confident tone, nice formatting — without its knowledge, producing models that sound right while being wrong. Depth of coverage, not surface polish, is what transfers real skill.

2. Curated generation: manufacture the data you wish existed. The phi “textbooks” approach — generate pedagogically ideal explanations and exercises, diversity-seeded so the corpus doesn’t collapse into repetition. Here the model isn’t smarter than the target; it’s a tireless writer executing a human-designed editorial vision. The leverage is in the curation loop wrapped around generation: seeding for diversity, filtering for quality, decontaminating against benchmarks.

3. Self-improvement with a filter. The most conceptually interesting kind: a model generates many attempts at hard problems, a verifier keeps only the successes, and the model trains on its own filtered wins. Math with checkable answers, code against unit tests, proofs through a proof checker — this loop (self-taught reasoner methods, rejection-sampling fine-tuning, and the RL-with-verifiable-rewards paradigm behind modern reasoning models) genuinely creates capability rather than copying it. Sampling plus verification acts as a search process: the model occasionally stumbles onto reasoning better than its average, the filter catches those moments, and training makes them the new average. AlphaZero pioneered exactly this shape — self-play generating positions, game outcomes verifying them — which is why RL people find nothing weird about “training on your own outputs.” The absolutely load-bearing part is the verifier. Without a reliable signal separating good from bad generations, the loop amplifies noise instead of skill.

4. Augmentation and privacy surrogates. The quieter, older kind: paraphrases, translations, simulated tabular records that mimic real distributions without exposing real people. Useful, established, less revolutionary.

Notice what the successful kinds share: an external source of signal — a stronger teacher, a human editorial standard, a verifier, a real distribution to mimic. Synthetic data works when generation is paired with judgment that doesn’t come from the generator itself. Which brings us to the famous failure mode.

Model collapse: the photocopy problem, and why the internet panicked

In 2023–24, “model collapse” papers hit mainstream news: train a model on data from a model, train the next on the next’s output, and quality degrades generation over generation — rare knowledge vanishes first (the distribution’s tails), then diversity, until outputs converge toward repetitive mush. The mechanism is real and intuitive: every generation is a sample from the previous model, samples under-represent the tails, and errors compound. Headlines concluded the internet was being poisoned by AI text and future models were doomed to eat their own decay.

The panic outran the result. The collapse experiments largely modeled a worst case: each generation training only on the previous generation’s output, discarding all original data — replacement, not accumulation. Follow-up work showed that when synthetic data accumulates alongside real data (which is what actually happens on the real internet and in real labs), degradation flattens or vanishes. And practitioner reality points the other way entirely: the best small models of the last two years are the most synthetic-heavy, not the least — because their synthetic data is curated, filtered, verified, and mixed with real text, not blindly recycled.

So my honest synthesis: model collapse is a true theorem about a pipeline nobody competent runs, and a false prophecy about the pipelines people actually run. The genuine risks live elsewhere — un-curated AI sludge polluting scraped web data (a contamination problem for everyone downstream), homogenization when many models learn from the same few teachers (a monoculture of style and blind spots), and teacher biases silently inherited at scale.

The discipline of doing it well

If I compress everything above into working rules, it’s these: generation is cheap; signal is the whole game — budget accordingly for verifiers, filters, and curation, because unfiltered generation is negative-value. Diversity must be engineered — seed generations from taxonomies, personas, documents, difficulty ladders, or the corpus quietly collapses into the model’s favorite phrasings long before any formal “model collapse.” Mix, never replace — synthetic tokens ride alongside real ones. Decontaminate — generated data has a nasty habit of paraphrasing benchmarks the generator memorized. And audit what the teacher smuggled in — style, politics, blind spots, and error patterns transfer just as efficiently as skills do.

The deepest way I’ve found to think about it: synthetic data converts compute into data — you spend inference FLOPs, plus a source of judgment, and receive training tokens. Seen that way, the scarcity story inverts. Human text is finite; compute keeps growing; verification (tests, checkers, reward models, principles) is improvable engineering. Which suggests the strange endpoint we’re already drifting toward: data stops being something models are given and becomes something models make — with humans supplying ever less of the text and ever more of the judgment about what counts as good. That division of labor — machines generate, standards select — might be the most important design pattern in modern ML. It’s also, I notice, roughly how evolution works.

One article left in this series: after all this theory about corpora and generation, the practical one — what actually goes into fine-tuning datasets, and how to build your own without stepping on the rakes.