When I started learning ML, I assumed the hard part was the model — architectures, optimizers, all the math I’ve written about so far. The longer I look at how frontier models are actually built, the more I believe something the practitioners keep repeating: past a certain point, the data is the model. Two teams with the same architecture and compute will get wildly different models depending on what they feed them. Yet the data pipeline is the least glamorous, least published part of the field — often literally a trade secret. This article is my attempt to map it: how a slice of the internet becomes a training corpus.
The raw material: what “trained on the internet” actually means
The starting point for nearly every open pretraining corpus is Common Crawl — a nonprofit that has been snapshotting the web since 2008, releasing petabytes of crawled pages. When people say a model was “trained on the internet,” they mostly mean “on a heavily processed subset of Common Crawl,” plus curated additions: code (GitHub), reference (Wikipedia), books, academic papers, forums (Reddit, Stack Exchange), and increasingly licensed and proprietary sources.
Here’s the number that reframed everything for me: raw Common Crawl is overwhelmingly garbage for training purposes. Boilerplate navigation, cookie banners, SEO spam, auto-generated pages, porn, malware listings, the same press release syndicated ten thousand times, pages that are 90% ads wrapped around a sentence of content. Serious pipelines discard the large majority of raw crawl — the famous corpora that sound huge (The Pile at 800GB, RefinedWeb, FineWeb at 15 trillion tokens) are what’s left after throwing most of the internet away. Pretraining data work is not collection. It’s demolition, done carefully.
The pipeline that does it has become fairly standardized, and each stage is more interesting than it sounds.
Stage 1: extraction — finding the text inside the page
A web page is HTML soup: markup, scripts, menus, footers, comment spam. Extraction pulls out the main content — and the choice of extractor measurably changes final model quality. Strip too aggressively and you lose structure (lists, tables, code blocks); too loosely and every document is polluted with “Accept cookies” and “Related articles.” Tools like trafilatura became standard because they thread this needle. It sounds like a detail; the FineWeb team showed extraction choices alone shift downstream benchmark scores. In data work, there are no details.
Stage 2: filtering — deciding what deserves to exist
Filtering happens in layers, cheapest first:
Language identification routes documents by language (and silently shapes which languages a model will be good at — low-resource languages often get crowded out here, a genuine equity issue in the field).
Heuristic quality filters are embarrassingly simple rules that do enormous work: drop documents that are too short or too long, with too few alphabetic characters, too many repeated lines, bizarre symbol ratios, no terminal punctuation, or a curse-word density suggesting spam. Gopher’s published rule list became a de facto standard. None of these rules “understand” quality — they’re proxies — but stacked together they remove oceans of junk for near-zero compute.
Model-based quality filtering is where philosophy enters. Train a lightweight classifier to score “quality,” and keep the high scorers. But quality by what definition? The GPT-3-era answer: “resembles books and Wikipedia.” That choice — high-brow reference English as the gold standard — quietly decided what the resulting models sound like and know. Modern pipelines (FineWeb-Edu is the clearest example) shifted the target: use an LLM to rate documents for educational value, distill those ratings into a fast classifier, keep the instructive slice. The result is one of the field’s most important recent lessons — filtering for educational content dramatically improves reasoning and knowledge benchmarks at the same token count. What you define as “good” becomes what the model is. There is no neutral filter.
Safety and compliance filtering removes toxic content, PII (emails, phone numbers get scrubbed or masked), and — increasingly — content whose owners opted out. This layer is where law, ethics, and engineering collide, and it’s in active motion: copyright litigation and robots.txt opt-outs are reshaping what future corpora may legally contain.
Stage 3: deduplication — the stage nobody skips twice
The web repeats itself absurdly: mirrors, syndication, boilerplate, quotes of quotes. Deduplication removes exact copies (via hashing) and near-duplicates (via MinHash/LSH — sketching documents so that mostly-overlapping ones can be found among billions without comparing every pair).
Why it matters so much: repetition is a multiplier on memorization. A document seen once contributes patterns; a document seen ten thousand times gets memorized verbatim — which wastes capacity, degrades generalization (well-documented: heavy duplication measurably hurts language modeling quality), and creates privacy and copyright hazards, since memorized text can be regurgitated. Dedup also prevents a subtler failure: benchmark contamination. If test sets leak into training data — and they do, constantly, because benchmarks get discussed and reposted across the web — evaluations silently turn into memory tests. Serious pipelines explicitly decontaminate against known evals, and contamination remains one of the field’s chronic credibility problems.
Stage 4: mixing — the recipe
After cleaning comes the most consequential set of dials: the mixture. What fraction code? (Even for non-coding models, code measurably helps reasoning — one of the field’s happier accidents.) How much multilingual text? Books vs. web? Academic papers? And how many epochs of each — high-value sources like Wikipedia are deliberately repeated a few times, while bulk web text is seen once (research on data-constrained scaling suggests roughly four epochs of good data is nearly as valuable as fresh data, and past that returns decay fast).
Then comes ordering. Naively you shuffle everything uniformly. Increasingly, labs run a curriculum: mid-training (annealing) phases that up-weight highest-quality and domain-targeted data (math, code, long documents) in the final stretch of pretraining, when the model is best positioned to absorb it. The published ablations here are sparse — mixtures are the crown jewels — but every open effort that documents its recipe (The Pile, Dolma, FineWeb, RedPajama) confirms the same thing: mixture changes rival architecture changes in their effect on final capability.
The pipeline is the moat
Step back and the strategic picture is stark. Architectures are published within months. Training code is open source. Compute is purchasable. The data pipeline — extraction choices, filter definitions, classifier training sets, dedup thresholds, the mixture, the curriculum — is where labs genuinely differ, and it’s precisely the part that papers describe in a vague paragraph. When an open effort like FineWeb publishes its full recipe with ablations, it’s a bigger gift to the field than most model releases.
And the frontier of this pipeline is moving toward a strange place: the web is finite, the high-quality slice of it more so, and models now consume tokens at a rate the internet doesn’t produce them. Estimates vary on when purely human-written public text effectively runs out for frontier training, but the direction is agreed. The responses — licensing private corpora, mining video and audio transcripts, and above all generating data with models themselves — are reshaping the pipeline I just described. Synthetic data is a big enough topic (and a big enough controversy) that it gets its own article, two posts from now.
Next up, though: the question that haunts every stage of this pipeline. Given a fixed budget, do you want more data, or better data? The answer changed twice in five years, and it’s one of the best stories in modern ML.