Of all the ideas in machine learning, this is the one that genuinely changed how I see the field: neural networks turn meaning into geometry. Words, images, songs, users, products — anything — can be represented as a point in a high-dimensional space, arranged so that similarity of meaning becomes proximity in space. This idea is called an embedding, and it is quietly the most deployed concept in all of AI: it powers search engines, recommendation systems, RAG pipelines, deduplication, clustering, and the internal workings of every LLM. This post is my attempt to explain it deeply.
The problem: computers need numbers, but not arbitrary ones
A model can’t operate on the word “cat.” It needs numbers. The naive encoding — assign each word an ID, or a one-hot vector with a single 1 in the “cat” position — is technically numeric but semantically dead: in one-hot space, “cat” is exactly as far from “kitten” as it is from “carburetor.” Every word is equally unrelated to every other. All the structure of language is thrown away before the model even starts.
An embedding fixes this by assigning each word a dense vector — a few hundred to a few thousand real numbers — and, crucially, by learning those numbers rather than assigning them. The training objective forces vectors to arrange themselves so that the geometry mirrors the semantics: “cat” ends up near “kitten,” “dog” nearby, “carburetor” far away in some other district of the space.
Where the meaning comes from: you are the company you keep
The engine behind all of this is the distributional hypothesis, a linguistics idea from the 1950s: words that occur in similar contexts have similar meanings. You can characterize “coffee” almost entirely by its neighbors — brewed, morning, cup, caffeine, bitter. Any word sharing those neighbors (“tea,” mostly) must mean something similar.
Word2vec (2013) turned this into a training objective of beautiful simplicity: given a word, predict its neighboring words (or the reverse). To succeed at this prediction game, the network is forced to give similar vectors to words with similar contexts — similarity isn’t programmed in, it’s the only way to win the game. The famous result was that the learned space had linear structure nobody explicitly asked for: vector(“king”) − vector(“man”) + vector(“woman”) ≈ vector(“queen”). The direction from “man” to “woman” encodes gender-ness; the same offset connects Paris→France with Berlin→Germany (capital-ness), and walk→walked with swim→swam (past-tense-ness). Concepts had become directions. Analogy had become arithmetic.
Two caveats worth internalizing. First, the arithmetic is cleaner in the famous examples than in general — it’s a real phenomenon, but a leaky one. Second, and more important: the space learns whatever co-occurs in the data, including its biases. The direction connecting man→woman also, in raw web-trained embeddings, connects programmer→homemaker. The geometry is a mirror of the corpus, not of truth — a fact with real consequences for every downstream system.
From word vectors to contextual vectors
Word2vec’s limitation: one vector per word, forever. But “bank” in “river bank” and “bank account” are different meanings, and a single point can’t be in two neighborhoods at once. The Transformer era resolved this: in models like BERT and every modern LLM, a token’s vector is computed fresh in context — each attention layer lets “bank” absorb information from “river” or “account” and drift to the appropriate region of the space. Embeddings stopped being a dictionary and became a computation.
This reframes what an LLM is doing internally, and it’s the mental model I now use: the residual stream — the vector flowing through the layers at each position — is an evolving embedding of “the meaning of the text so far, as relevant to what comes next.” Layer by layer, attention moves information between positions and MLPs refine each one, sculpting a point in representation space; the final prediction is read off that point’s geometry. Interpretability research strengthens the picture: many human-interpretable features genuinely correspond to directions in these spaces (the linear representation hypothesis), and techniques like sparse autoencoders are used to un-mix the thousands of overlapping features superimposed in each vector. “Meaning as geometry” isn’t a metaphor about neural nets. It’s fairly literally how they work.
Sentence and document embeddings: the workhorse of applied AI
The most commercially important descendant of these ideas is the embedding model: a network trained to map an entire sentence, paragraph, or document to a single vector such that semantically related texts land close together. This is trained with contrastive learning — pull matched pairs (a question and its answer, a query and a relevant document, two paraphrases) together in the space, push mismatched pairs apart. Similarity is then just a cosine between vectors.
This one capability is the foundation of:
- Semantic search — embed the query, embed the documents, return nearest neighbors. Matches “how do I fix a flat” to “repairing a punctured tire” with zero shared keywords, something lexical search structurally cannot do.
- RAG (retrieval-augmented generation) — the same search, wired into an LLM: retrieve the most relevant chunks from your knowledge base, stuff them into the prompt, ground the model’s answer in them. The retrieval half of nearly every serious LLM application is embeddings.
- Recommendation — embed users and items in a shared space (from behavior: things interacted-with by similar users drift together); recommending is nearest-neighbor lookup around the user’s point.
- Clustering, dedup, anomaly detection, classification — once everything is a point in meaning-space, decades of geometric algorithms apply to raw text and images directly.
Making nearest-neighbor search fast over billions of vectors is its own engineering field — approximate nearest neighbor (ANN) indexes like HNSW graphs and IVF/product-quantization trade a sliver of recall for orders of magnitude in speed, and “vector database” became a product category on the back of it.
And embeddings aren’t confined to language: models like CLIP train an image encoder and a text encoder into a shared space, where a photo of a dog sits near the sentence “a photo of a dog.” Cross-modal geometry is what makes text-to-image search — and much of multimodal AI — work.
What high-dimensional space is actually like
A last idea that took me a while, because human intuition is built for three dimensions and embedding spaces have thousands. High-dimensional space is roomy in a specific, useful way: the number of directions that are nearly-orthogonal to each other grows exponentially with dimension. That means a 1,000-dimensional space can host vastly more than 1,000 distinguishable concepts, each as its own quasi-independent direction, interfering with the others only slightly — the phenomenon interpretability researchers call superposition. It’s why a single vector of a few thousand numbers can simultaneously encode a token’s topic, tone, syntax, language, position in an argument, and a thousand other features at once. The curse of dimensionality, viewed from this angle, is a blessing: meaning needs room, and high dimensions provide it.
That’s the whole idea, and I find it genuinely beautiful. Intelligence — at least the machine kind — runs on a map where near means alike, concepts are directions, and understanding a thing means placing it well among everything else. Every search box, every recommendation, every RAG pipeline, every LLM forward pass is navigation on that map. Learning to think in embeddings is learning to think in the native language of modern AI.