A model learns by adjusting weights, and weights only know how to do one thing: multiply and add numbers. So how do you hand a system like that something as human as a word, a photograph, or a song, and expect anything useful to come out the other side?
Every part of a model, from its very first step to its last, depends on a working answer to that question. The answer is the vector.
What a vector actually is
A vector is an ordered list of numbers where each position has an agreed meaning.
Picture describing a person on a form: height in one box, age in another, income in a third. You wouldn't dream of filling in age where height belongs, because the position of each number does real work. A vector works the same way. `[3, 7, 1]` is a vector. So is a list of a thousand numbers. What turns a random pile of numbers into a vector is that agreement about position: entry one always means the same kind of thing, whichever vector you're looking at.
Early attempts at language processing tried to keep words as words: matching exact strings, counting how often "bank" appeared near "river" versus "money," building huge lookup tables by hand. That approach breaks the moment a computer has to generalize, to notice that "happy" and "joyful" mean roughly the same thing even though they share no letter. Researchers needed a format where similarity could be measured arithmetically instead of listed by hand, and a fixed-length list of numbers is exactly that format.
Who decides what each position in the vector should stand for?
Nobody does, at least not by hand. During training the model organizes its own positions however proves useful for predicting the next word. One position might end up loosely tracking formality, another something closer to sentiment, and plenty will settle into patterns with no clean English name at all. That isn't a flaw. It's what lets the format scale to millions of words and concepts without a human sitting down to define each one.
The Nook of Wonder Theorems & Beautiful Patterns — notation you'll actually see
A vector with n numbers is said to live in n-dimensional space, written:
v ∈ ℝn
Read as: "v is a vector of n real numbers." A 5-number vector is in ℝ⁵. Real models commonly use vectors with hundreds or thousands of entries, and part 4 of this series explains why that scale turns out to be necessary rather than wasteful.
Two vectors of the same length can be added or scaled, entry by entry, with no mixing across positions:
[1, 2, 3] + [0.5, -1, 2] = [1.5, 1, 5]
2 × [1, 2, 3] = [2, 4, 6]
This is the same operation behind the weighted sum from part 1. A weighted sum scales a vector by weights, position by position, and adds the results into a single number.
Seeing it work: movie taste as coordinates
Once something is described by numbers, “how similar are two things” becomes “how close are two points.”
Say you describe your movie taste using two numbers: how much you like action, and how much you like comedy, each on a scale from 0 to 1. That pair is a vector, and it plots cleanly on a simple graph.
Look at what happened without any extra machinery. “Similar taste” became “nearby on the graph.” “Different taste” became “far apart on the graph.” No paragraph describing anyone’s preferences was needed. Two numbers and ordinary geometry did the job, in a form a computer can measure rather than read.
Real models don’t stop at two numbers describing taste in movies. They use hundreds or thousands of positions, each one a dimension the model discovered was worth tracking during training: tone, formality, sentiment, topic, grammatical role, and a long tail of dimensions with no tidy human label at all. A word, a sentence, an image, all of them end up as a single point sitting somewhere inside that much larger space. The two-number movie example is a toy, but the underlying operation, plotting meaning as position, is what a full-scale model does at every layer.
Not only words
It’s tempting to assume this trick is specific to language, so it’s worth naming plainly that it isn’t. An image becomes a vector by describing patterns in pixels instead of words. A song becomes a vector by describing patterns in sound. A shopper’s history can become a vector describing buying habits. The input changes, but the model’s task stays constant: find a useful list of numbers that captures what matters, and quietly drop what doesn’t.
what makes vectors genuinely powerful
Doing arithmetic on meaning
Because vectors are just numbers, ordinary addition and subtraction on them can produce results that line up with real relationships between ideas.
Word vectors from an early model, word2vec, gave the clearest demonstration of this.word2vecthe 2013 Google model that made learned word vectors widely available Take four ordinary words and ask whether one relationship among them matches another:
king − man + woman = queen
Nobody instructed the model that “queen” is what results from swapping the gender of “king.” That relationship fell out of the geometry on its own. The direction you travel from “man” to “woman” closely matches the direction from “king” to “queen,” because that direction came to encode something like gender as a side effect of learning from enormous amounts of text.
This mattered for more than being a neat trick. It was strong early evidence that these vector spaces weren’t convenient storage. They were capturing real structure in how concepts relate, well enough that basic arithmetic could navigate it. That finding is a large part of why embeddings, the next article, became a central technique rather than a curiosity.
The result doesn’t hold for every pair of words you might try, and modern models represent meaning in richer, more context-sensitive ways than this picture suggests. The core idea still holds: direction in vector space can carry meaning, not just position.
The Nook of Wonder Theorems & Beautiful Patterns — why subtraction reveals direction
Subtracting one vector from another gives you the vector that would move you from the second point to the first: literally, the direction and distance between them.
woman − man = Δ (a direction representing "shift toward feminine")
Adding that same Δ to a different word applies the same shift starting from a different point:
king + Δ ≈ queen
In practice this gets checked by computing king − man + woman, then searching the model's vocabulary for whichever real word's vector sits closest to that resulting point, using cosine similarity, covered next. "Queen" usually wins that search, and not because it was hard-coded anywhere. It is the nearest real word to the computed point.
measuring "close" properly
Two different ways to measure “similar”
Vectors let you measure similarity two ways, straight-line distance or shared direction, and picking the right one matters.
Ask a plain computer program whether “happy” and “joyful” are similar words, using only their spelling. It has nothing to work with. The letters barely overlap, and a naive string comparison calls them unrelated. Symbols alone don’t carry meaning a machine can act on. Vectors fix it by turning “how similar” into “how close,” which geometry can answer.
As raw text
"happy" and "joyful" share no letters in matching positions. Spelling alone gives no path to comparing meaning.
"happy" ≠ "joyful" (as strings)
As vectors
Both land as nearby points in the learned space, measurably close, regardless of spelling.
distance("happy", "joyful") ≈ small
If two points are close together, is “close” the same as “pointing the same direction”? Those are two different measurements, and each answers a slightly different question.
The first, straight-line distance, is the ruler measurement from school geometry extended past two dimensions. The second, the angle between two vectors, ignores how far out each point sits and asks only whether they’re oriented the same way. For comparing meaning, the second tends to win, for a plain reason: a short movie review and a long one can express nearly identical opinions while landing at very different distances from the origin, because one has more words pushing its vector further out. Measuring angle instead of distance cancels that out.
The Nook of Wonder Theorems & Beautiful Patterns — Euclidean distance vs. cosine similarity, worked out
Euclidean distance is the straight-line ruler measurement between two points, the Pythagorean idea from school geometry extended past two dimensions:
distance(a, b) = √[ (a₁−b₁)² + (a₂−b₂)² + … + (aₙ−bₙ)² ]
Worked example with the movie-taste vectors a = [0.9, 0.3] (you) and b = [0.2, 0.85] (Friend B):
distance = √[(0.9−0.2)² + (0.3−0.85)²]
= √[(0.7)² + (−0.55)²] = √[0.49 + 0.3025] = √0.7925 ≈ 0.89
Smaller means closer. This measurement is sensitive to magnitude, how far out a point sits overall, which is not always what you want to be measuring.
Cosine similarity avoids that by comparing only the angle between two vectors, ignoring their length entirely:
cos(θ) = (a · b) / (‖a‖ ‖b‖)
The top, a · b, is the dot product: multiply matching positions together, then add them. a·b = a₁b₁ + a₂b₂ + … + aₙbₙ. The bottom divides by each vector's length (its norm, ‖a‖), rescaling the result to always land between −1 and 1, no matter how long either vector is.
Worked example, comparing you to Friend A instead:
a = [0.9, 0.3] b = [0.8, 0.4]
a · b = (0.9×0.8) + (0.3×0.4) = 0.72 + 0.12 = 0.84
‖a‖ = √(0.9² + 0.3²) ≈ 0.949 ‖b‖ = √(0.8² + 0.4²) ≈ 0.894
cos(θ) = 0.84 / (0.949 × 0.894) ≈ 0.99
0.99 sits close to the maximum of 1, meaning these two vectors point in nearly the same direction: very similar taste, and in a language model, very similar meaning. A score near 0 signals unrelated directions; near −1 signals close to opposite. Because it cancels out magnitude, cosine similarity is the default choice for comparing meaning across most modern AI systems, including the embeddings covered next.
where this shows up
- Search: Finding results by meaning, not just matching keywords
- Recommendations: “More like this” means nearby in vector space
- Every model input: Text, images, and audio all get converted to vectors first
- Analogy tools: “This is to that as X is to ___” solvers, built on vector arithmetic
- Fraud detection: Flagging transactions whose vectors sit far from a user’s usual pattern
- Deduplication: Spotting near-identical documents or images by vector distance