Every training run begins with a budget question: you have finite compute and finite data — what maximizes capability? For a while the field’s answer was simply “more.” Then “more, but allocated correctly.” Then, in one of the most interesting reversals in recent ML, “less, but immaculate” started winning benchmarks. This article traces that argument — Chinchilla, LIMA, and the phi models are its landmarks — because I think the quality-vs-quantity question is the single best lens for understanding why today’s small models embarrass yesterday’s giants.
Round one: quantity wins (and everyone was still doing it wrong)
The scaling-laws era (which I covered two series ago) established that loss falls predictably as you add parameters, data, and compute. The first practical takeaway, from Kaplan et al., skewed toward model size — grow parameters fast, data slower. GPT-3 embodied it: 175B parameters, ~300B tokens.
Chinchilla (2022) corrected the ratio: for compute-optimal training, parameters and tokens should grow together — roughly 20 tokens per parameter. GPT-3-era models were massively undertrained; DeepMind’s 70B Chinchilla, fed 1.4T tokens on the same compute as the 280B Gopher, beat it decisively. Then inference economics pushed further: since serving cost scales with model size, it pays to train small models far past their “optimal” token count — which is exactly the Llama-and-descendants playbook, 8B-class models drinking 15T+ tokens.
Notice what all of round one shares: every token is treated as identical. Chinchilla math counts tokens the way physics counts mass — quantity, no character. The next two results attacked exactly that assumption.
Round two: LIMA — a thousand examples against fifty thousand
LIMA (“Less Is More for Alignment,” Meta 2023) asked how much fine-tuning data it takes to turn a pretrained base model into a capable assistant. The prevailing assumption: a lot — tens of thousands to millions of instruction examples, plus RLHF on top.
LIMA’s team fine-tuned Llama-65B on exactly 1,000 examples — no RLHF, no reward model — but those thousand were obsessively curated: carefully selected community answers from Stack Exchange and Reddit plus hand-written demonstrations, filtered for quality, diversity of task, and a uniform response style. In human evaluations, this thousand-example model produced responses competitive with — sometimes preferred over — models trained on vastly more data with far more elaborate pipelines.
The paper’s explanation is the part worth internalizing, the superficial alignment hypothesis: almost everything the model knows — facts, reasoning, language — was already learned in pretraining. Fine-tuning doesn’t teach knowledge; it teaches which subdistribution of its abilities to present: the format, persona, and style of being an assistant. And selecting a style is a low-information task — a thousand consistent examples specify it fine, whereas fifty thousand inconsistent ones specify it badly, teaching the model to average over sloppy and careful answers alike.
The practical corollaries changed how everyone I read now builds fine-tuning sets: in SFT, a mediocre example is not neutral — it is damage. Diversity of tasks matters more than volume. Consistency of style matters more than either. And an important boundary: LIMA is a claim about alignment/style, not about knowledge or reasoning — you cannot LIMA your way into capabilities pretraining didn’t provide, and later work confirmed harder skills (math, code) do benefit from much more, and more rigorous, fine-tuning data. Less is more for the things that are actually low-information.
Round three: phi — “Textbooks Are All You Need”
The most radical attack on token-equality came from Microsoft Research’s phi series (2023). The provocation of the first paper is right in the title: Textbooks Are All You Need. Their argument: web-scale corpora are mostly noise — repetitive, poorly explained, error-ridden — and a token of pedagogically excellent text (clear explanations, worked examples, progressive difficulty: textbook-nature) is worth many tokens of sludge. So they built small corpora of “textbook-quality” data — aggressively filtered web content selected for educational value, plus synthetic textbook-style text and exercises generated by a stronger LLM — and trained tiny models on them.
phi-1, at 1.3B parameters trained on ~7B tokens, hit coding benchmarks that models an order of magnitude larger struggled with. phi-2 (2.7B) traded blows with 25× larger models on reasoning. The series continued scaling the philosophy, and the message landed across the industry: data quality doesn’t just shift the scaling curve — it can substitute for a startling amount of scale. In Chinchilla’s terms: quality changes the constant in the scaling law, and the constant turns out to be enormous.
Fair criticisms exist and are worth carrying: benchmark-adjacent synthetic data raises contamination-flavored questions (are we teaching the distribution of the test?); phi models were long noted to be benchmark-brilliant and somewhat brittle in open-ended use; and “textbook quality” inherits the biases of whichever model wrote the textbooks. But the core result survived scrutiny and reshaped practice — the clearest public confirmation being FineWeb-Edu: filter a general web corpus with an educational-value classifier, train on the surviving slice, and watch reasoning benchmarks jump at identical token counts. Every serious lab now runs quality classifiers over pretraining data. The phi bet became the industry default.
The synthesis: quality is leverage on quantity
So who wins the argument? The unsatisfying, correct answer: it was never either/or — the results compose into something like a hierarchy of data value:
Scale still rules the floor. Pretraining needs trillions of tokens; no curation makes 10B tokens produce a frontier generalist. Chinchilla’s allocation logic remains the budget backbone.
Quality multiplies every token. Filtering, dedup, and educational selection shift the whole curve — the same compute buys visibly more capability. This is now the cheapest known way to “increase” your compute.
Curation dominates the top of the funnel. By fine-tuning, examples number in the thousands and each one carries real weight; here, quality isn’t a multiplier — it’s nearly the whole game, per LIMA.
The pattern connecting all three: the later in training a token arrives, the more its quality matters. A trillion pretraining tokens can absorb noise by sheer averaging; a thousand SFT examples cannot. Which is also why the modern curriculum practice — saving the best data for mid-training and annealing phases — makes sense: it moves quality to where quality counts most.
There’s one more force pressing on this debate, and it decides where the field goes next: the internet is running low. High-quality human text is being consumed faster than humanity writes it, which means quantity — the resource round one assumed infinite — is becoming the binding constraint. The escape hatch everyone reached for is the one phi already opened: if great data is scarce, manufacture it. Whether that works, when it collapses, and why “model-generated data” went from taboo to standard practice in about two years — that’s the next article.