The Bit Explainers · Building Blocks · Part 4

Dimensions: why AI thinks in thousands of directions at once

Every entry in an embedding is a separate axis of meaning. Here's what that actually buys a model, and what it costs.

Good Omens Studio12 min readMachine Learning

Last time, we turned words into dense vectors, a few hundred numbers instead of one arbitrary ID, arranged so that similar things land close together. What we skipped over is what each of those numbers is actually doing. A vector with 300 entries has 300 dimensions, and each one is, in principle, an independent axis a concept can vary along. That’s a strange claim if you’ve only ever thought about dimensions as up-down, left-right, and forward-back. This article is about what a dimension really means once you leave physical space behind, why models end up needing hundreds or thousands of them, and why “more dimensions” isn’t automatically better.

What a dimension actually is

In everyday use, “dimension” means one of the three directions you can move in physical space, plus maybe time as a fourth if you’re being fancy. That’s a special case of a much more general idea, and it’s the general version that matters here. A dimension, mathematically, is just an independent way something can vary, something you could change without necessarily changing anything else about it. Height, weight, and age are three dimensions you could use to describe a person, and none of them is secretly a rearrangement of the other two. You need all three numbers to pin the person down; two isn’t enough, and a fourth wouldn’t hurt but also wouldn’t be required.

An embedding vector works the same way, just with far more axes and none of them labeled in advance. A 300-dimensional word embedding isn’t describing height, weight, and age. It’s describing 300 separate, learned axes of variation that turned out to be useful for predicting context, whatever those axes end up meaning.

In one sentence

A dimension in an embedding is one independent axis a concept can vary along, and a model with hundreds of dimensions is tracking hundreds of such axes at once, not one physical direction repeated many times.

Intuition

Imagine describing a restaurant to a friend using only three numbers: price, spiciness, and distance from home. Three dimensions, and a lot of restaurants would already separate out nicely along them. But you'd quickly want more: how noisy it is, how fast the service is, whether it takes reservations, how good the dessert menu is. Each new question you could reasonably ask, where the answer is independent of the others, is another dimension. A model faced with the entire space of human language ends up with an enormous list of such questions, and it needs a slot for each one.

Technical

Formally, an n-dimensional vector space is the set of all possible lists of n numbers, and a dimension is one coordinate axis of that space. Nothing requires n to be 3 or even a number you could visualize. The math of vectors, addition, scaling, dot products, cosine similarity, works identically whether n is 3 or 3,000. A model's embedding dimension is how many independent coordinates it has allotted to represent every word, and that number is a design choice made before training starts, usually between 100 and 4,000.

But if we can only picture three dimensions, how can a thousand-dimensional space mean anything at all?

It means something the same way a spreadsheet with a thousand columns means something: you never have to picture the whole thing at once. You only need the arithmetic to keep working consistently, and addition, distance, and angle are defined the same way whether there are 3 columns or 3,000. Visualizing it is a convenience for humans, not a requirement for the math.

Why one model needs so many axes

Language is stubbornly multi-dimensional in the literal sense above. Two words can be close on one axis of meaning and far apart on another, simultaneously, and a good representation has to hold both truths at once. “Hot” and “cold” are opposite in temperature but similar in that they’re both temperature words, both adjectives, both capable of describing weather or food or tempers. A single-number “temperature scale” for words would flatten all of that into one line and lose almost everything else. Give the model a second axis for whether a word is literal or figurative, a third for formal or casual register, a fourth for whether it’s used mostly in food contexts or mostly in emotional ones, and each additional axis lets it keep two words close on the traits they share while still telling them apart on the ones they don’t.

Real embedding dimensions rarely line up with single human concepts this neatly. That’s a simplification for intuition, not a description of what you’d find inspecting dimension 47 of a trained model. The underlying pressure is real regardless: the more independent ways two things can be similar or different, the more axes a model needs before it runs out of room to keep those distinctions separate.

There’s a further wrinkle, since it comes back later. A model doesn’t need one dimension per concept. Two or more unrelated ideas can share the same handful of dimensions, as long as they don’t tend to show up in the same examples at the same time. A dimension that tracks “past tense” in most sentences might quietly also help represent something about formality in the rare sentences where tense isn’t doing much work. This overlapping, called superposition, is part of how models pack far more distinctions into a vector than the raw dimension count suggests.superpositionseveral unrelated features sharing one dimension, since they rarely co-occur The cost is that any single dimension becomes hard to read on its own.

The curse of dimensionality

More axes sounds like a strictly good thing, and up to a point it is. Past that point, high-dimensional spaces start behaving in ways that work against you, a set of effects collectively nicknamed the curse of dimensionality. The clearest version of it: as you add dimensions, the space gets so vast, so quickly, that the data you have gets spread impossibly thin across it. Points that were meaningfully close together in a 10-dimensional space can end up looking almost equally far apart from everything once you’re in 10,000 dimensions, because there’s so much room that “close” stops being a useful distinction. Nearest-neighbor comparisons, the exact tool an embedding space is built for, get less reliable the more room there is to spread out in.

This is part of why embedding dimensions have a practical ceiling rather than climbing forever. A bigger vector can represent more distinctions, but it also needs proportionally more training data to actually learn good values for every one of those extra coordinates, and it makes every downstream similarity comparison a little noisier if the added dimensions aren’t pulling their weight.

None of this makes high-dimensional embeddings unusable in practice. Systems built on them are running everywhere already. It means engineers work around the effect rather than ignore it: approximate nearest-neighbor search algorithms, the kind that power large-scale semantic search, are built to stay fast and accurate as dimensionality climbs.approximate nearest-neighbor searchANN: trading a little accuracy for speed by not checking every vector Dimensionality reduction is also often applied before a search rather than after, trimming a vector down to the handful of axes carrying the most signal.

The Nook of Wonder Theorems & Beautiful Patterns — why more room can mean less signal

Take random points inside an n-dimensional unit cube. The ratio between the farthest and nearest neighbor distances, for a fixed point, tends toward this as n grows large:

(distmax − distmin) / distmin → 0 as n → ∞

In plain terms: the farthest point and the nearest point start looking almost equally distant. Concretely, for 2 random points in a 2D square, distances vary a lot, some pairs are close, some are far. Run the same experiment with 1,000 random dimensions instead of 2, and nearly every pair of points ends up clustered around roughly the same distance from each other.

2D: distances spread across a wide range
1,000D: distances cluster near a single value

That's the curse in one line: "nearest neighbor" becomes a much weaker signal once there's too much empty room for the data to actually fill.

low dimensions many dimensions
In low dimensions, points sit at meaningfully different distances. Add enough dimensions and nearly every point ends up about the same distance from every other.

Choosing dimensionality in practice

Because of that tradeoff, dimensionality is one of the first real decisions made when a model is designed, and different systems land in very different places on purpose.

Fewer dimensions

Faster to store, faster to compare, cheaper to search across millions of items. Good when the distinctions that matter are relatively coarse, or when speed and memory matter more than capturing every nuance.

More dimensions

Room to keep fine-grained distinctions separate, better at subtle relationships. Costs more to store and search, needs more training data to fill well, and risks the curse of dimensionality if pushed too far past what the data can support.

System Typical dimensions Why
Classic word2vec embeddings ~100–300 Small vocabulary task, speed mattered more than nuance
BERT-style sentence embeddings ~768 Needs to capture sentence-level context, not just single words
Large language model internal representations 1,000s–10,000s+ Tracking grammar, facts, tone, and reasoning state simultaneously

There’s no universally correct number. It’s a balance struck against the specific task, the amount of training data available, and how much compute is on hand to search across the resulting vectors later.

It’s also not a decision made once and forgotten. Inside a large language model, the dimensionality of its internal representation, often called the hidden size, ripples into nearly every other design choice: how many attention heads it can run in parallel, how large its intermediate layers need to be, how much memory each token costs to hold in context. A model’s dimension count isn’t just a fact about its embeddings; it’s closer to a foundation the rest of the architecture gets built on top of, which is part of why we’ll come back to it once we get to how transformers actually work.

What we do when there’s too much to look at

Humans still need to inspect these spaces sometimes: to debug a model, to sanity-check that similar things are actually landing near each other, to build a chart for a presentation. Nobody can look at a 1,000-dimensional scatter plot, so the standard move is dimensionality reduction, techniques like PCA or t-SNE that squash a high-dimensional space down to two or three axes while preserving as much of the original structure as possible.PCAprincipal component analysis, which rotates a space to keep its largest axes of variance
t-SNEa nonlinear projection that keeps local neighbourhoods intact, for plotting
It’s lossy, since some real distinctions get flattened away, but it’s usually enough to see whether clusters that should be separate actually are.

Which raises the question of whether dimension 212 of some embedding means something specific, the way a spreadsheet column means “age.” Mostly no. Individual dimensions in a trained embedding rarely map cleanly onto single human concepts. Much of what a model learns gets spread across many dimensions at once, and single dimensions often encode blends of several ideas rather than one clean label. The two-axis temperature-word example earlier was a teaching simplification. Reduction techniques don’t recover clean labels either; they find whichever handful of directions happen to capture the most variation for a particular chart.

  • A 2D scatter plot of "similar customers" clustering together in a dashboard
  • Genetic or medical research visualizations that compress thousands of measurements into one chart
  • "Similar songs" maps in music apps, built by projecting a high-dimensional taste space down to something plottable
  • Compression tools that store a rough approximation of something using far fewer numbers than the original
  • Not a system where each axis has a plain-English label ready to read off