The Bit Explainers · Building Blocks · Closing

You Built the Whole Chain

Article 1 opened with a single weight, multiplying a single number. Nothing since has introduced anything more magical than that. Here's what stacking it, honestly, article by article, actually bought you.

Good Omens Studio6 min readMachine Learning

Ask someone who’s only heard the one-line summary, a large language model just predicts the next word, to explain why that prediction takes a warehouse of specialized hardware. They’ll shrug. Ask someone who just finished twelve articles built on a single weighted sum, and the shrug turns into an answer with names attached: attention, backpropagation, quantization, PagedAttention.PagedAttentionthe vLLM technique for storing the KV cache in pages instead of one contiguous block Nothing new appeared out of nowhere. Every one of the last twelve articles was built for the one before it.

Retracing the chain, in one breath

A weight multiplied a number. Enough weights, arranged deliberately, became a vector. A vector, positioned by training, became an embedding that placed a word’s meaning in space, which only made sense because that space had enough dimensions to hold it. Batches of those vectors traveled together as tensors, and tokens turned out to be the unit those vectors were representing all along. Weighted sums stacked into layers, useless on their own until an activation function gave them a fold worth stacking. Attention let one token’s fold borrow context from another’s. Training and inference turned out to be the same forward pass doing two different jobs, one carrying a backward pass the other never touches. That backward pass needed a loss function to aim at and gradient descent to act on the aim. The resulting model still had to be sized, quantized, routed, and served to millions of people without falling over, which is where the series left off.

Say it that fast and it sounds like a lot happened. The same handful of moves, weighting, folding, comparing, correcting, got reapplied at a slightly larger scale each time. That repetition is the whole trick, and it’s the same trick the model itself is built on.

Before and after

What actually moved in your head

Being precise about what changed matters here, since “I understand AI better now” is exactly the vague claim this series has argued against since article 1. What changed is narrower: specific phrases that used to be a wall now have a mechanism behind them.

Before this series

"Attention," "gradient descent," and "quantization" were words you'd nod along to, correctly associated with AI, without a mechanism attached to any of them.

After this series

Each of those words points to something you could sketch on a napkin: a query compared against a key, a step taken against a slope, a weight stored in fewer bits.

What you can now actually do

The tests worth trying on yourself

A working mental model shows up as an ability, not a memory, so the honest way to check whether this series actually landed is to try explaining a few things out loud, from scratch, without flipping back to an earlier article. If any of these feel shaky, that’s not a failure, it’s just a pointer back to exactly which article to reread.

  • Explain why a stack of layers with no activation function is secretly just one layer in disguise
  • Walk through what a query, a key, and a value each actually do in one attention step
  • Say why training needs several times more memory than the same model needs to answer a question
  • Explain what a gradient is pointing toward, and why subtracting it makes a model less wrong
  • Say why a bigger model is reliably slower per token, independent of how good the hardware is
  • If an explanation you give still leans on "it just does," that's the one worth revisiting first

Where the series goes from here

Building Blocks was the foundation, not the ceiling

Two series are complete now: Foundations, which mapped out what a large language model is and where it sits in the wider AI conversation, and Building Blocks, which just spent twelve articles building the actual mechanism underneath it, weight by weight, layer by layer. Together they cover the what and the how. What’s queued next covers the refinements: reinforcement learning, fine-tuning, LoRA, and quantization examined properly rather than summarized in a paragraph the way article 11 had to.

Every one of those upcoming topics is, not coincidentally, a direct extension of a mechanism already sitting in this series. Fine-tuning and LoRA are gradient descent from article 10, aimed at a small set of a model’s weights instead of all of them. Reinforcement learning is the loss function from article 10 replaced with a very different kind of signal, one built from feedback about a whole response rather than a single next token. Quantization already got a full technical treatment in article 11 and will get an even closer look. None of it will ask you to unlearn anything from these twelve articles. It’s the same foundation, extended.

Reinforcement learningA new kind of loss signal

Training a model against feedback on whole responses, not just the next token.

Fine-tuning & LoRAGradient descent, aimed narrower

Adjusting a small slice of an already-trained model's weights for a specific task.

Quantization, in depthArticle 11's shortcut, unpacked

A full technical look at how fewer bits per weight actually gets implemented.

Twelve articles, four numbers each

  • 1 weight: Where the entire series started, and the unit every later mechanism ultimately reduces back down to.
  • 2 jobs: Training and inference, the same forward pass, carrying completely different responsibilities.
  • 3 letters: Q, K, and V, the three projections that let a token ask a question and get an answer from its neighbors.
  • 12 articles: The full chain, weight to real world, each one standing on exactly the one before it.