The Bit Explainers · Building Blocks · Introduction

Twelve Pieces, One Working Model

Somewhere between "it's just predicting the next word" and a hundred-billion-dollar training run sits an actual mechanism. This series builds it, piece by piece, starting from the smallest one there is.

Good Omens Studio7 min readMachine Learning

Here’s a sentence you’ve probably heard, maybe said yourself: a large language model just predicts the next word. It’s technically true, and it explains almost nothing.

If next-word prediction were the whole story, a pocket calculator’s worth of logic should be able to do it. It can’t, and the gap between that one-line summary and what’s actually running underneath is what this series is for. An actual walk through the machinery, one layer at a time, until “predicts the next word” stops being a magic trick and becomes a specific, buildable thing.

Why the one-line summary undersells everything

If you read the Foundations series, this site’s first run at explaining large language models, you already know roughly what an LLM is, how it fits into the wider AI landscape, and where the real debates about it sit. That series answered the “what.” This one answers the “how,” from the ground up: one mechanism at a time, each built directly on the mechanism before it rather than skipping ahead to the exciting parts.

That ground-up approach matters more than it sounds. Almost every popular explanation of how these models work either stays so high-level it could describe a search engine just as well, or jumps into a research paper’s notation and loses anyone not already fluent in linear algebra. Both leave the same gap: you can repeat the phrase “next-token prediction” but you can’t explain why a weighted sum needs an activation function, why attention needs three separate projections instead of one, or why the same forward pass costs wildly different amounts depending on whether it’s training or answering a question. This series closes that gap by never letting a later article lean on a concept it hasn’t already earned.

The shape of the climb

Twelve articles, three honest acts

Building Blocks is one continuous dependency chain that falls into three acts, and knowing the shape of the climb ahead makes each article land better. The first act gathers the raw ingredients every later mechanism needs: a single weight, a vector, an embedding, the dimensions a vector lives in, the tensors that carry batches of them, and a fresh look at tokens now that there’s real machinery to describe what happens to them. Nothing in that act is a neural network yet. It’s the vocabulary the rest of the series speaks.

The second act builds the machine: how simple weighted sums stack into layers, how attention lets a token borrow context from its neighbors instead of reading in isolation, and what separates training a model from running one someone else already trained. The third act gets practical: the mechanics of how a model learns from its mistakes, the tradeoffs behind choosing a model’s size and serving it to real users, and a close look at what all of this is doing in the world right now.

Parts 1–6The ingredients

Weights, vectors, embeddings, dimensions, tensors, tokens. The raw vocabulary everything else in the series speaks.

Parts 7–9The machine

Layers that fold, attention that reads context, and the split between training a model and running one.

Parts 10–12The discipline

How a model learns from being wrong, how it gets sized and served, and where all of it stands right now.

1234567 weightsattention
The series is a literal stack. Article 12 only makes sense standing on top of article 1, not beside it.

The full roadmap

Where each piece fits

You don’t have to read this series in one sitting, but it is built to be read in order. Article 8’s explanation of attention leans directly on article 3’s embeddings and article 7’s layers without re-explaining either. Article 11’s memory footprint assumes you’ve seen article 9’s forward and backward pass. Skipping ahead will mostly make sense, since the prose tries to stand on its own, but it won’t build the same intuition.

# Article What it adds to the stack
1 Weights & Parameters The single number every other mechanism eventually reduces to
2 Vectors Bundling many weights into one meaningful direction
3 Embeddings Turning words into vectors that capture meaning
4 Dimensions What it means for a vector to have hundreds of entries
5 Tensors Carrying whole batches of vectors through a model at once
6 Tokens, Revisited Rebuilding tokenization now that vectors and dimensions exist
7 Neural Network Layers Stacking weighted sums with a fold that actually bends them
8 Inside the Transformer Letting tokens read context from each other, mechanically
9 Training vs. Inference Why learning and answering are two very different jobs
10 Loss Functions & Gradient Descent How a model finds out it was wrong, and by how much
11 Small Models, Big Models & vLLM Sizing a model and serving it to real, simultaneous users
12 Real World AI Every mechanism above, checked against what’s shipping right now

How to actually read it

What you’ll run into along the way

Every article follows the same handful of recurring moves, so the format gets out of the way and the content can do the work. Concepts with real depth get broken into three passes, one plain sentence, one intuitive analogy, one precise mechanism, in that order, so you can stop reading at whichever level answers your question. Mathematical moments get worked out by hand in a dedicated aside rather than asserted, with small enough numbers to follow every step. And every article closes by naming the takeaway plainly, no dressed-up ending.

  • Each article assumes everything before it, and nothing after it
  • Worked math examples use small, hand-checkable numbers, never a wall of notation
  • Every new concept gets a plain-language version before the technical one
  • This series won't teach you to build a model from scratch: it teaches you to understand what one is doing

What you'll be able to do

  • Explain a fold: Say exactly why stacking layers without an activation function buys a model nothing at all.
  • Trace an attention weight: Walk through how a query, a key, and a softmax turn into one token borrowing meaning from another.
  • Explain a hosting bill: Point to the exact mechanism, memory-bandwidth-bound generation, that makes a bigger model cost more per token.
  • Question a headline: Read an AI adoption statistic and know which part is mechanism and which part is still an open question.