In 2017, a paper with the almost arrogant title “Attention Is All You Need” replaced decades of accumulated architecture research with one mechanism. Every major language model today — and increasingly every major vision, audio, and biology model — is a Transformer. Understanding this one architecture means understanding the skeleton of modern AI. Here’s the mental model that finally made it click for me.
The problem: sequences resist parallelism
Before Transformers, sequence models were recurrent: RNNs and LSTMs read text one token at a time, threading a running “memory” vector through the sequence. Two fatal flaws. First, information from early tokens had to survive being squeezed through that single vector across hundreds of steps — long-range dependencies faded. Second, and commercially decisive: step 500 can’t be computed before step 499, so recurrent training is fundamentally sequential, and GPUs — machines built for doing thousands of things at once — sat half idle.
The Transformer’s founding bet: throw away recurrence entirely. Process every token in the sequence simultaneously, and let a mechanism called attention handle the relationships between them. All positions computed in parallel, perfectly matched to GPU hardware. That single property is why models could scale from millions to trillions of tokens of training data.
Tokens become vectors
A Transformer doesn’t see words. Text is first chopped into tokens — subword chunks from a fixed vocabulary (typically 32k–200k entries), so that even rare words decompose into known pieces. Each token ID looks up a learned embedding: a vector of maybe a few thousand numbers. These vectors are learned during training, and they end up encoding meaning as geometry — similar tokens sit near each other, and directions in the space correspond to semantic relationships.
One catch: since we abandoned recurrence, the model has no inherent sense of word order — attention by itself treats the input as a bag of vectors. So positional information is injected explicitly, either added to the embeddings or (in most modern models) baked into the attention computation via rotary position embeddings (RoPE), which encode positions as rotations and generalize better to long contexts.
Attention: every token asks questions of every other token
Here is the heart of it, and the metaphor that made it real for me: attention is a soft database lookup.
From each token’s vector, the model computes three things (each via a learned matrix):
- a query — “here’s what I’m looking for,”
- a key — “here’s what I contain, for matching purposes,”
- a value — “here’s the actual content I’ll hand over if you match me.”
Every token’s query is compared (dot product) against every other token’s key, producing a relevance score for each pair. The scores are scaled, then softmaxed into weights that sum to 1 — a probability distribution over the sequence. Each token then updates itself with the weighted sum of everyone’s values.
Concretely: in “The animal didn’t cross the street because it was too tired,” the token “it” emits a query that — in a trained model — matches the key of “animal” far more strongly than “street.” So “it” absorbs the value of “animal,” and from that layer onward, the vector at position “it” contains the information that it refers to the animal. Ambiguity resolved, by learned lookup, in one parallel step. No information had to crawl token-by-token across the sentence; every pair of positions is one hop apart.
Three refinements complete the picture:
Multi-head attention. Instead of one lookup, run 8–128 in parallel, each with its own learned query/key/value projections attending in a lower-dimensional subspace. Different heads learn different relationship types — interpretability work has found heads tracking syntax, coreference, even induction heads that implement “this pattern happened before; copy what came next,” a mechanism strongly linked to in-context learning.
Causal masking. A language model predicting the next token must not peek ahead. During training, attention scores to future positions are set to −∞ before the softmax, so every position attends only backward. This is what lets a single forward pass over a document train the model on every next-token prediction in it simultaneously.
The cost. Every token attending to every token is quadratic in sequence length — the central scaling pain of the architecture, and the reason for a whole research industry: FlashAttention (exact attention, computed memory-efficiently), grouped-query attention (share keys/values across heads to shrink the KV cache), sliding-window and hybrid schemes for long contexts.
The MLP: where the knowledge lives
Attention gets the fame, but it’s only half of each Transformer block. After attention moves information between positions, a feed-forward network (MLP) processes each position independently: expand the vector to ~4× its width, apply a nonlinearity, project back down. Two-thirds or more of a model’s parameters live in these MLPs, and interpretability research increasingly views them as the model’s key-value memory — where facts and associations are stored. A useful division of labor: attention routes; the MLP computes and remembers.
Each block wraps both parts with two pieces of plumbing that made depth possible at all: residual connections (each sublayer’s output is added to its input, so every layer refines a running “stream” of information rather than replacing it, and gradients flow cleanly through hundreds of layers) and normalization (keeping activation scales stable; modern models normalize before each sublayer, which trains far more stably).
Then you stack. A block’s job: attend, mix, refine. A model is 30–100+ of these blocks. Early layers handle surface patterns and syntax; middle layers, the model’s semantic workhorse; late layers, task-specific assembly of the actual prediction. At the very top, a final projection maps the last-layer vector to a score for every token in the vocabulary — softmax those, and you have the next-token distribution.
Three body plans, one organ
The original paper described an encoder-decoder for translation. The field has since settled into three configurations of the same components: encoder-only (BERT-style — every token attends in both directions; ideal for understanding tasks like classification and embeddings), decoder-only (GPT-style — causal masking, trained to predict the next token; this is essentially every modern LLM), and encoder-decoder (T5, Whisper — still the natural shape when input and output are genuinely different sequences). The decoder-only design won the era for a simple reason: one objective (predict the next token), applied to any text, at any scale, subsumes almost every task — as long as the task can be phrased as text continuation. Which, it turns out, nearly everything can.
Why this architecture, of all architectures, won
I keep coming back to three properties. It’s parallel, so it converts hardware into capability more efficiently than anything before it. It’s general — attention makes no assumptions about grids, adjacency, or even modality, which is why the same architecture now handles images (ViT), audio, protein structure (AlphaFold’s Evoformer is attention at heart), and multimodal everything. It scales smoothly — where older architectures saturated, Transformers keep converting more data and compute into predictably lower loss, the empirical fact that scaling laws (next article’s topic) made precise.
There’s a humbling meta-lesson in that. The Transformer isn’t the most biologically plausible design, nor the most theoretically elegant. It won because it fit the hardware and kept improving with scale. In deep learning, the architecture that best rides the compute curve beats the architecture that best encodes our ideas about intelligence — a pattern the field calls the bitter lesson, and one I now see everywhere.