The Bit Explainers · Building Blocks · Part 3

Embeddings: where similar ideas learn to live near each other

Vectors gave every idea an address. Embeddings are what happens when a model chooses those addresses on purpose, so that meaning has geography.

Good Omens Studio12 min readMachine Learning

Last time, a vector was a list of numbers standing in for a point, or a direction, in some space. On its own that’s an empty trick. A list like [0.2, -1.4, 0.9] means nothing until something decides what each position stands for and where things should land relative to each other.

Embeddings are the part of the system that makes that decision. They’re the reason a vector for “puppy” ends up closer to “dog” than to “umbrella.”

Why numbers need a neighborhood

Picture the crudest possible way to turn words into numbers: give every word in the dictionary an ID. “Cat” is word 4,001. “Dog” is word 4,002. “Umbrella” is word 19,884. This works, in the sense that a computer can now store and retrieve words as numbers. But it throws away everything that actually matters. Word 4,001 is no more mathematically related to word 4,002 than it is to word 19,884, even though cats and dogs share almost everything in common and neither shares much with an umbrella. The numbering is arbitrary, so any math you do on it is meaningless. You can’t average two IDs and get something sensible, and you can’t measure how “close” two words are, because closeness in ID-space has nothing to do with closeness in meaning.

What you actually want is a numbering scheme where the arithmetic itself carries information: where nearby numbers mean related things, and where the distance between two points is a real, usable signal. That’s the whole problem embeddings solve.

There’s a name for the arbitrary-ID scheme, and it’s worth knowing because embeddings are usually introduced as its replacement: a one-hot vector. Instead of a single ID number, “cat” becomes a vector the length of the entire vocabulary, all zeros except for a single 1 sitting in the “cat” position. That fixes the “arbitrary number” problem, since every word is now equally distant from every other word by design, but it introduces a new one: the vectors are enormous, one entry per word in the language, so tens of thousands of dimensions, and almost entirely empty. An embedding compresses that sparse, meaningless vector into something dense, maybe a few hundred numbers instead of tens of thousands, where every entry is doing real work and the distance between any two vectors actually means something.

In one sentence

An embedding is a vector assigned to something, a word, an image, a product, a user, so that things which are alike end up as vectors that sit close together in space.

Intuition

Think about how a well-run library is organized. It isn't alphabetical, it's arranged by subject, so a book on soil chemistry sits a few shelves from a book on composting, even though their titles share no letters. Someone browsing the gardening section stumbles into related ideas just by walking a few steps. An embedding does the same thing to numbers: it arranges them by what they mean, not by an arbitrary ID, so that "browsing nearby" in vector space actually surfaces related things.

Technical

An embedding layer is a big table with one row per possible input, every word in a vocabulary say, and a fixed number of columns, often a few hundred to a few thousand.embedding layerthe lookup table that turns a token ID into a vector Each row starts out as random numbers. As a model trains on real data, those rows get nudged by gradient descent until rows for things that behave similarly in the training data end up numerically similar. Nobody hand-writes the coordinates; training discovers them.

But if every row starts as random noise, how does a model know which words are supposed to end up near each other?

It doesn't know in advance. It figures it out from a much simpler signal: which words tend to show up in similar company.

Learning meaning from context

Here’s the idea that makes the whole thing work, and it’s older than deep learning by decades: a word’s meaning can be approximated by the company it keeps. “Cat” and “dog” show up in a lot of the same sentences, next to words like “vet,” “leash,” “food,” and “sleeping.” “Umbrella” shows up next to “rain” and “forgot” instead. You never have to tell a model what a cat is. You show it enough sentences and ask it to get good at a narrow, mechanical task: given a word, guess which words are likely to appear around it. Or the reverse, given the surrounding words, guess the missing one in the middle.

That guessing task is deliberately boring. The model doesn’t care about grammar rules or dictionary definitions. It only cares about minimizing prediction error, over and over, across millions of sentences. But to get good at that boring task, it’s forced to develop a useful side effect: words that tend to appear in the same contexts have to end up with similar vectors, because the model is reusing the same internal machinery to predict around all of them. Nobody designed that outcome directly. It falls out of the training objective as the cheapest way to get the predictions right.

“Context” here means a small window of neighboring words, often five or ten on either side of the target. That window size turns out to matter. A narrow window tends to group words that are near-synonyms and interchangeable in a sentence, like “happy” and “glad.” A wider window pulls in looser, topical relationships, like “happy” drifting toward “celebration” or “wedding.” Neither is more correct; they’re tuned for different jobs, and picking one is a decision someone makes before training starts.

dog cat kitten rain umbrella storm manwoman kingqueen
A simplified 2D slice of a real embedding space. Related words cluster, and the man-to-woman direction runs roughly parallel to the king-to-queen one.

Measuring closeness: distance and direction

Once words are vectors, “similar” becomes something you can compute. The most common measure isn’t the straight-line distance you’d first guess at. It’s the angle between two vectors, called cosine similarity: two vectors pointing in almost the same direction score close to 1, vectors at a right angle score close to 0, vectors pointing opposite ways close to -1. Direction carries the meaning here, because a model that’s more confident about a word (further from the origin) shouldn’t count as “further away” in meaning from a less confident one.

The striking part is that directions start to mean something consistent, not just positions. In a well-trained word embedding space, the direction you’d travel to get from “man” to “woman” is roughly the same direction you’d travel to get from “king” to “queen.” The model was never told anything about gender or royalty as concepts. It picked up a consistent “male to female” direction purely because that pattern kept showing up across enough sentences that encoding it was the most efficient way to make good predictions.

The Nook of Wonder Theorems & Beautiful Patterns — how close is close?

Cosine similarity between two vectors a and b:

cos(a, b) = (a · b) / (‖a‖ ‖b‖)

Say "cat" = (4, 1) and "kitten" = (3, 2), two toy 2D embeddings.

a · b = (4)(3) + (1)(2) = 14

‖a‖ = √(4² + 1²) = √17 ≈ 4.12
‖b‖ = √(3² + 2²) = √13 ≈ 3.61

cos(a, b) = 14 / (4.12 × 3.61) ≈ 0.94

Now compare "cat" = (4, 1) against "umbrella" = (-2, 5):

a · b = (4)(-2) + (1)(5) = -3

cos(a, b) = -3 / (4.12 × 5.39) ≈ -0.13

0.94 is close to 1 (nearly the same direction, so very similar). -0.13 is close to 0 (barely related). Real embeddings run this exact calculation across hundreds of dimensions instead of two, but the arithmetic is identical.

Beyond single words

Nothing about this idea is specific to words. Any input a model deals with can be embedded, as long as you can define a training objective where “things that behave alike should end up numerically alike.” Images get embedded so that two photos of golden retrievers land near each other and far from a photo of a bicycle. Whole sentences and paragraphs get embedded so that “the store closed early” and “the shop shut down ahead of schedule” land close together despite sharing almost no words at all. Products, songs, and even individual users get embedded in recommendation systems, so that “people who behave like this user” is a distance you can measure rather than a rule you have to write by hand.

Word embeddings

One fixed vector per word, learned once from surrounding context. "Bank" gets a single vector, blending its river-bank and money-bank senses together, since the model never sees the word in isolation from a sentence.

Sentence & document embeddings

One vector for a whole chunk of text, built by combining its words in context. "The bank was flooded" and "the bank raised interest rates" get pushed toward different regions of space, because the surrounding words disambiguate the meaning before the vector is finalized.

That second column is closer to what modern systems actually use. A single word rarely carries enough information on its own, but a sentence embedding can capture “this whole passage is about a river overflowing its banks,” which is exactly the kind of thing a search engine or a retrieval system needs to compare against a query.

The same trick extends further than text. An image embedding model is trained on a different task, often matching photos to the captions people actually wrote for them, but the underlying goal is identical: push visually and semantically related images close together, push unrelated ones apart. A recommendation system does the same thing with a user’s viewing or purchase history instead of a sentence, embedding “everything this account has clicked on” into a single vector, so that finding what to recommend next becomes the same nearest-neighbor search as finding a related word. The training signal changes with the domain. The underlying geometry doesn’t.

What this makes possible

Once meaning is a location in space, a lot of previously hard problems turn into nearest-neighbor lookups. Want documents related to a question, even when they share no keyword with it? Embed the question, embed every document, return whichever land closest. Want to catch a scam email that never uses the words your filter was trained on? Embed it and check whether it lands near the cluster your filter already treats as suspicious.

This is also the mechanism behind retrieval-augmented generation, where a language model searches a knowledge base by embedding similarity before it writes an answer, instead of relying purely on what it memorized during training.retrieval-augmented generationRAG: searching a document collection before writing, so answers can cite sources

The tradeoff is that an embedding is only as good as the data and the objective it was trained on. A model trained mostly on news articles will build a space where “bank” leans financial by default. One trained on outdoor recreation content might lean the vector toward rivers. Embeddings don’t discover a universal structure of meaning. They discover whatever structure was consistently useful for predicting the specific data they were shown, which is a narrower and more defensible claim than “the model understands language.”

  • Search engines returning results that match your intent, not just your exact words
  • "Customers who bought this also bought…" recommendations on shopping sites
  • Music and video apps building a playlist around a "vibe" rather than a genre tag
  • A chatbot pulling the right paragraph out of a document before it answers your question
  • Not a system that looks up dictionary definitions, it never stored any