Everything this series has covered so far, weights, vectors, embeddings, dimensions, tensors, assumes the input is already a tidy list of numbers. It never is, not at first. A model starts with raw text: sentences, typos, emoji, half-finished words, three different languages in one paragraph. Somewhere between “the user typed this” and “the model computed that,” text has to be broken into a fixed set of reusable pieces and handed a number apiece. That first cut is called tokenization, and it’s worth a second, deeper look now that you know what a vector and an embedding actually are, since the choices made at this very first step ripple through everything downstream.
Splitting text into pieces
The obvious first guess is to split on whitespace and call each word a token. It works for a sentence or two, but it breaks down fast at scale. Every language has more distinct words than any model could reasonably keep a slot for; English alone has hundreds of thousands, before you count names, typos, made-up internet words, and words borrowed from other languages. A model built around whole words either needs an enormous vocabulary to cover them all, or it hits words it’s never seen and has no way to represent them at all, a dead end usually called the out-of-vocabulary problem.
The fix that stuck, used by essentially every modern language model, is to split text into pieces smaller than a whole word but usually bigger than a single letter, called subword tokens. Common whole words like “the” or “dog” often survive as a single token, because they’re frequent enough to earn their own slot. Rarer or longer words get broken into meaningful fragments instead, “unbelievable” might become “un,” “believ,” and “able.” The model never runs out of vocabulary this way, because any text, however unfamiliar, can always be broken down further, all the way to individual letters or bytes if it has to.
Whitespace splitting has a second problem beyond vocabulary size, one that doesn’t show up until you leave English: not every language actually separates words with spaces. Written Chinese and Japanese, for instance, run words together without spaces at all, so a tokenizer built around “split on whitespace” has nothing to split on there. A subword tokenizer sidesteps this too, since its merge rules are learned from whatever characters and byte patterns actually appear in the training text, regardless of whether that language happens to use spaces as word boundaries.
A token is the smallest chunk of text a model treats as a single unit of input, and it's usually a piece of a word rather than a whole word or a single letter.
Think about how a shipping company handles packages of wildly different sizes using a small, fixed set of standard box sizes. They don't keep a custom box on hand for every possible item; they break oversized shipments down into a handful of standard boxes instead, and something small enough already fits in one box on its own. Tokenization does the same thing to language: instead of needing a custom slot for every possible word, it keeps a fixed, manageable set of "boxes," and anything that doesn't fit as a whole word gets broken into pieces that do.
The algorithm behind most modern tokenizers is byte-pair encoding, or a close relative of it.byte-pair encodingBPE: building a vocabulary by repeatedly merging the most frequent adjacent pair It builds its vocabulary of subword pieces by statistics rather than by grammar. It looks at a huge amount of text, finds whichever pair of adjacent symbols shows up together most often, and merges that pair into a new single unit. Repeat that tens of thousands of times, starting from individual characters, and the pieces that survive as their own tokens are the ones that showed up often enough in real text to be worth a dedicated slot. "ing" and "tion" earn tokens fast; rare letter combinations stay split apart.
But if the pieces are chosen by frequency rather than meaning, do they actually line up with anything a person would recognize as a meaningful chunk?
Sometimes, though not always. Byte-pair encoding has no idea what a prefix or a syllable is; it only tracks which symbol pairs appear together often. It rediscovers a lot of real linguistic structure, prefixes, suffixes, common word stems, purely because those patterns are frequent. It will just as happily produce a split that looks arbitrary to a human reader, because frequency was the only thing it was optimizing for.
Watching the merges happen
The training process behind byte-pair encoding is short enough to trace by hand on a toy example, which makes it one of the few algorithms in this series you can watch work rather than take on faith. The same repeat-count-and-merge loop, run at a much larger scale on a much larger body of text, produced the vocabulary inside every major tokenizer in use today.
The Nook of Wonder Theorems & Beautiful Patterns — merging a tiny vocabulary by hand
Start with a made-up mini corpus, already split into individual characters, with counts of how often each word appears:
"lo w" × 5
"lo w e r" × 2
"n e w e s t" × 6
"w i d e s t" × 3
Step 1: count every adjacent character pair across the corpus. The pair e s shows up in both "newest" (×6) and "widest" (×3), for 9 total, the most of any pair. Merge it into a single new symbol:
e + s → es
Step 2: recount pairs with "es" now treated as one unit. The pair es t appears 9 times (all the "newest" and "widest" occurrences), still the top count. Merge again:
es + t → est
Keep repeating: count all adjacent pairs, merge the most frequent one, recount, merge again. Stop once the vocabulary hits whatever target size was chosen ahead of time, often in the tens of thousands for a real model. "est" survives as a single token here because "newest" and "widest" made it common enough to earn one, exactly the kind of frequency-driven decision described above.
From tokens to the rest of the pipeline
A token by itself is still just a piece of text, so the next step is to give each unique token in the vocabulary a fixed ID number, the same arbitrary-numbering move discussed back in the embeddings article. Once the sentence has been chopped into tokens and each token swapped for its ID, those IDs get looked up in an embedding table, the exact learned lookup from Part 3, and each ID comes back as a dense vector positioned by meaning. Stack all those vectors for every token in a sentence, and for every sentence in a batch, and you get a tensor, the exact structure from Part 5. Text to tokens to IDs to embeddings to tensors: that’s the entire path from a sentence you typed to the numbers a model actually computes with.
One consequence of this ordering is worth flagging: a tokenizer and the model trained alongside it are locked together. The embedding table has exactly one row per token ID, so a model can only ever be as flexible as the vocabulary it was tokenized with, and swapping in a different tokenizer after training would leave every ID pointing at the wrong row, effectively scrambling the model’s entire vocabulary. This is why tokenizers are trained once, early, and then frozen for the rest of that model’s life; changing the very first step of the pipeline means retraining everything downstream of it.
Choosing a vocabulary size
How many distinct tokens to allow, the vocabulary size, is a real design decision with real tradeoffs, not a technical detail settled once and forgotten.
Smaller vocabulary
Fewer rows in the embedding table, cheaper to store and train. Costs more tokens per sentence on average, since more words have to be split into several pieces, which means longer sequences for the same amount of text.
Larger vocabulary
More whole words and phrases get their own single token, so sequences stay shorter for the same text. Costs more memory for the embedding table, and rarer tokens get less training exposure each, since the same amount of text is now spread across more distinct pieces.
| Tokenizer family | Typical vocabulary size |
|---|---|
| Early word2vec-style vocabularies | ~30,000–50,000 |
| BERT-style tokenizers | ~30,000 |
| Modern large language model tokenizers | ~100,000–250,000 |
The trend over time has mostly pointed toward larger vocabularies, since a shorter token sequence for the same sentence means less computation spent per unit of actual text, and compute is usually the scarcer resource. But it’s still a balance, not a straight line upward, since an oversized vocabulary spreads training signal thin across too many rarely used tokens, and the embedding table itself grows by one full row for every additional token, a real memory cost that adds up across hundreds of thousands of entries.
Where tokenization quietly shapes what you see
Knowing how tokenization actually works explains a handful of odd, otherwise-mysterious model behaviors. A model asked to count the letters in a word is really working from a small number of subword chunks, not individual letters, so counting anything smaller than a token is an awkward task for it, closer to asking a person to count individual grains of sand in a handful rather than count the handfuls. Text in languages that weren’t well represented in the tokenizer’s training data tends to split into far more tokens per word than English does, which matters directly for cost and speed in any system billed or measured by token count. Even something as small as a leading space can change which token ID a word maps to, since many tokenizers treat “cat” and “ cat“ as two entirely different entries in the vocabulary.
Numbers cause their own version of this problem. A tokenizer trained mostly on prose doesn’t necessarily split digits the same predictable way every time. “1234” might come out as one token in one context and as “12” plus “34” in another, depending on which digit groupings were common in its training data. That inconsistency is part of why arithmetic on multi-digit numbers is harder for a language model than it looks: it isn’t seeing a clean sequence of individual digits to carry and add the way long addition teaches, only whatever irregular chunks the tokenizer produced for that specific number.
None of these quirks are bugs waiting to be patched out. They’re the direct consequences of a design decision made around frequency and compute efficiency, with no built-in awareness of spelling, arithmetic, or fairness across languages. Recognizing that tokenization does this work, quietly, before anything else in the pipeline gets a turn, is usually enough to stop being surprised by the specific ways it shows its seams.
- API pricing measured in tokens rather than characters or words
- A chatbot struggling to count letters in an unusual word
- Autocomplete suggesting the rest of a word after only a few letters are typed
- Non-English text costing noticeably more tokens for the same sentence length
- Not the same thing as a "word count," even though the two numbers are often close