Most of machine learning research is messy and empirical — try something, see if the benchmark goes up. Then, in 2020, a team at OpenAI published a result with a completely different character. Model performance, they showed, follows smooth mathematical laws: predictable curves relating loss to model size, dataset size, and compute, holding across seven orders of magnitude. Suddenly ML had something that looked like physics — and the industry has been spending accordingly, in billions of dollars, ever since. This is the story of those laws, the famous correction to them, and what they actually mean.
The discovery: loss is a power law
Kaplan and colleagues trained families of Transformer language models, systematically varying three quantities: N (number of parameters), D (number of training tokens), and C (total compute). Their finding: when nothing else is the bottleneck, the loss on held-out text falls as a power law in each factor. Loss scales roughly as N^(−0.076) — multiply the model size by 10, and the loss drops by a fixed, predictable ratio. Same shape for data and for compute.
Three properties of this result made it revolutionary rather than merely interesting:
Smoothness. No plateaus, no walls, no diminishing-returns cliff within the measured range — just a clean straight line on a log-log plot, spanning models from thousands to billions of parameters.
Predictability. You can fit the curve on small, cheap models and extrapolate the loss of a model 1,000× larger before you spend the money to train it. This transformed frontier training from a gamble into an engineering projection — famously, GPT-4’s final loss was predicted accurately from experiments using roughly 1/10,000th of its compute.
Architecture insensitivity. Width vs. depth, attention heads, minor architectural choices — within the Transformer family, these barely bent the curve. What mattered was scale. This finding demoted architecture tinkering and promoted a new discipline: figuring out the optimal allocation of a compute budget.
That allocation question is where the story gets interesting.
Chinchilla: everyone was training models wrong
Kaplan’s paper included allocation guidance: as compute grows, spend most of it on model size, scaling parameters much faster than data. The field obeyed. GPT-3 was 175B parameters trained on ~300B tokens; other labs built even larger models on similar token counts.
In 2022, DeepMind’s Hoffmann et al. — the Chinchilla paper — redid the measurement more carefully. A subtle flaw in the original setup (chiefly, learning-rate schedules not tuned to each training length) had skewed the conclusion. The corrected law says: parameters and data should scale in equal proportion. Double the model, double the tokens. As a rule of thumb, a compute-optimal model wants roughly 20 tokens per parameter.
By that math, GPT-3 and its contemporaries were dramatically undertrained — huge brains, starved of reading. DeepMind proved the point directly: Chinchilla, a 70B model trained on 1.4T tokens, outperformed the 280B Gopher trained with the same total compute on ~4× fewer tokens. A model a quarter the size won, simply by allocating the identical budget correctly. It’s one of the cleanest experimental results in the field, and it reshaped every training run that followed.
The inference correction: why models got small again
Chinchilla optimizes one thing: best loss for a fixed training budget. But a deployed model is trained once and run billions of times — and inference cost scales with model size. If you’re going to serve a model at scale, it pays to train a smaller model far past its Chinchilla-optimal token count: you spend extra training compute to buy a cheaper model forever.
This is exactly what the industry did. Llama-class models pushed to hundreds of tokens per parameter — small models trained on 15T+ tokens, wildly “overtrained” by Chinchilla’s metric, and vastly more useful per dollar of serving cost. The loss curve bends with diminishing returns past the optimum, but it keeps going down, and the economics keep it worth it. The modern regime, in one line: Chinchilla tells you the optimum for training; inference economics tells you to overshoot it on data.
Emergence: smooth loss, jumpy abilities
Here’s the wrinkle that makes scaling laws philosophically strange. The loss improves smoothly — but specific capabilities often don’t. Arithmetic, multi-step reasoning, certain benchmarks: near-zero performance across model sizes, then a rapid climb past some scale threshold. These “emergent abilities” are why each model generation feels qualitatively different rather than incrementally better.
There’s a genuine debate about how real emergence is. One influential analysis argued much of it is a measurement artifact: score with a harsh all-or-nothing metric (exact-match on a 5-digit sum) and gradual underlying progress looks like a sudden jump; score with a smooth metric (per-digit accuracy, log-likelihood) and the curve smooths out. The truth seems to be both: underlying competence grows continuously with loss, but usefulness is often thresholded — a model that gets a multi-step task 90% right per step still fails the task most of the time, until per-step reliability crosses the point where whole chains succeed. Smooth physics, discontinuous experience.
Either way, the practical upshot stands: the headline loss number understates what scaling buys, because capabilities keep unlocking along the curve in ways the loss alone doesn’t advertise.
Scaling isn’t just pretraining anymore
The original laws describe one axis: pretraining compute. The last few years added others, each with its own returns curve:
- Data quality scaling. Chinchilla counts tokens as if all tokens are equal. They aren’t. Aggressive filtering, deduplication, and curriculum (and increasingly synthetic data) shift the entire curve — a smaller model on better data beats a bigger one on sludge. This is now a primary competitive axis, and arguably the real lesson of the “small model” era.
- Post-training. Instruction tuning and preference optimization convert raw predictive ability into usable behavior, at a tiny fraction of pretraining compute — enormous capability per FLOP.
- Inference-time scaling. Let the model think longer — extended chains of reasoning, search, self-verification — and performance on hard problems climbs with test-time compute, following scaling curves of its own. The 2024–25 reasoning-model wave established this as a third pillar: you can now buy capability at training time or at answer time, and trade between them.
The frontier question is no longer “do returns to scale continue?” — so far, stubbornly, they do — but “which scaling axis is cheapest for the next increment?” Raw web data is approaching exhaustion, which pressures the data-quality and synthetic-data axes; inference-time reasoning is expensive per query, which pressures distillation. The laws don’t tell you where capability comes from. They tell you something more useful: capability has a price, the price follows a curve, and whoever reads the curves best allocates their compute — and their money — where a dollar buys the most intelligence.
That’s the real legacy of scaling laws. They turned “how smart can we make it?” from a research mystery into a budgeting question. Whether that’s thrilling or unsettling probably depends on the day, but it’s the single most important fact about how modern AI actually gets built.