The Bit Explainers · Building Blocks · Part 5

Tensors: the shape data takes inside AI

Vectors were one row of numbers. Tensors are what you get when a model needs to stack rows into grids, and grids into stacks of grids.

Good Omens Studio11 min readMachine Learning

We’ve spent two articles treating a vector as the basic unit of meaning: a list of numbers, one per dimension, positioned in space. That covers a single word, a sentence, an image, one thing at a time. Real training runs never move data one thing at a time. A model processes thousands of sentences together, each made of many word vectors, in batches, across many layers, all at once.

The container that holds all of that simultaneously is called a tensor.tensorthe general name for an array of numbers with any number of axes It turns out to be the same idea as a vector, generalized one step further.

From lists, to grids, to stacks

Start from the bottom and build up. A single number, like 7, is the simplest case: no direction, nothing to index into, just a value. Line several numbers up in a row and you get a vector, which is what the last two articles were about. Stack several vectors of the same length on top of each other and you get a grid of numbers, rows and columns, better known as a matrix. A photograph in grayscale is naturally a matrix already: one number per pixel, arranged in rows and columns exactly the way the pixels sit on screen.

Keep going and the pattern doesn’t stop being useful. A color photograph needs three numbers per pixel instead of one, red, green, and blue, so you stack three matrices together, one per color channel. Now you need a third axis just to say which channel you’re looking at. Stack a batch of a hundred such photographs together for training, and you need a fourth axis for “which photo in the batch.” Each time you add a new independent way the data is organized, you add another axis, and a matrix isn’t equipped to have more than two.

It helps to notice that this isn’t a special AI trick bolted onto ordinary math, it’s the same generalization that vectors already represented one step further. A vector generalized a single number into a list; a tensor generalizes that list into however many axes the situation actually calls for. Nothing about addition, multiplication, or any other operation needs to be reinvented at each step. The rules that work on a matrix are the same rules working on a rank-4 tensor, just applied across more axes at once.

In one sentence

A tensor is a container for numbers arranged along any number of axes at once, a scalar has zero, a vector has one, a matrix has two, and a tensor is the general name for all of these plus anything with three or more.

Intuition

Think of a single spreadsheet as a matrix: rows and columns, two axes. Now imagine a whole workbook of spreadsheets, one tab per month. You'd need three coordinates to find any single number, which sheet, which row, which column, and that whole workbook is doing the job of a 3-axis tensor. Stack a shelf of such workbooks, one per year, and you're up to four axes without the underlying idea changing at all. Nothing new is happening conceptually past the jump from a single sheet to several; it's the same trick, repeated.

Technical

The number of axes a tensor has is called its rank, or sometimes its order (unrelated to the "order" of a matrix in other contexts). The size along each axis is called its shape, usually written as a tuple like (100, 224, 224, 3) for a batch of 100 color images that are 224 pixels tall, 224 wide, with 3 color channels. Every deep learning framework, PyTorch and TensorFlow both take their names directly from this data structure, represents data internally as tensors, and every operation a neural network performs, multiplying, adding, transforming, is defined to work across however many axes the tensor happens to have.

But why not just use several separate vectors or matrices instead of bundling everything into one object with extra axes?

Because keeping the structure explicit is what lets hardware process it efficiently. A tensor with a batch axis tells the GPU "these hundred images are independent, work on all of them in parallel," in a way that a hundred separate matrix variables never could. The shape isn't bookkeeping. It's an instruction for how the computation should actually be organized.

Reading a shape like a sentence

Once you get used to it, a tensor’s shape tells a small story about what the data actually is, in the order the axes are listed. A shape of (32, 128) for a batch of sentence embeddings reads as “32 sentences, each represented by a 128-number vector.” A shape of (16, 3, 224, 224) for a batch of images reads as “16 images, 3 color channels each, 224 pixels tall, 224 pixels wide.” Add a time axis and a shape like (8, 30, 3, 224, 224) becomes “8 video clips, 30 frames each, in color, at 224 by 224 resolution.” Nothing about the underlying numbers changed conceptually between these examples. Only the number and meaning of the axes did.

The order those axes are listed in isn’t arbitrary. It’s a convention the whole surrounding codebase has to agree on, and mixing conventions is a common source of bugs. Some frameworks put the batch axis first, others put channels before height and width, and a tensor with the right shape but the wrong axis order produces nonsense rather than an obvious error. The arithmetic still runs; it just runs on the wrong interpretation of what each number means.

batch channel height width which item which channel which row which column
Each added axis is one more independent question you can answer about a number's position.

Why deep learning needed this specific idea

None of this would matter much if it were only a bookkeeping convenience. The reason tensors became the foundational data structure of the field is hardware. GPUs perform the same simple operation, usually multiplication and addition, across enormous numbers of values at once, rather than one value at a time the way an ordinary processor would. A tensor’s shape is exactly the information a GPU needs to split a huge computation into independent chunks it can run in parallel: process every image in the batch axis at once, every channel at once, every pixel at once, and combine the results according to well-defined rules.

This is also why batching exists as a concept at all. Feeding a model one example at a time wastes almost all of a GPU’s capacity, since most of its parallel processing units would sit idle waiting for the next example. Bundle a hundred examples into a single tensor with a batch axis instead, and the same hardware that struggled with one example now processes a hundred in roughly the same amount of time, because the extra axis is precisely what tells the chip “these are independent, spread the work across everything you’ve got.”

A regular CPU is built the opposite way: excellent at doing one complicated thing quickly, in strict order, but with only a handful of cores available to split work across. A GPU trades that flexibility for raw count, thousands of much simpler cores, each capable of doing the same small operation on a different piece of data at the same instant. Tensor operations are designed specifically to hand a GPU exactly that kind of work: one identical instruction, applied across thousands of independent slots defined by the tensor’s shape. Choose the wrong shape, say, cramming everything into a single giant vector instead of organizing it into proper batch and channel axes, and the GPU has no clean way to split the work, so a lot of that parallel hardware ends up sitting unused.

Reshaping without losing anything

Data frequently needs to move between shapes as it flows through a model, and two operations handle almost all of it: reshaping and broadcasting. Reshaping takes a tensor’s existing numbers and rearranges which axis each one belongs to, without changing how many numbers there are or what any individual value is. Flattening a batch of 224-by-224 images into a batch of 50,176-long vectors, for instance, is a reshape: same total count of numbers, just relabeled from a 2D grid per image into a 1D list per image. Broadcasting is different: it’s a set of rules that let a smaller tensor act as if it were stretched to match a larger one during an operation, without actually copying any data, so a single bias vector can be added to every row of a matrix without a duplicate copy needing to exist for each row.

Both operations exist because copying data is expensive, and a model that copied a full tensor every time it needed to combine two differently shaped ones would waste enormous amounts of memory and time doing nothing but duplication. Reshaping avoids that by never touching the underlying numbers at all, only their labels. Broadcasting avoids it by pretending the smaller tensor is bigger than it is, on the fly, for exactly as long as the calculation needs it and no longer.

Reshaping

Rearranges the same numbers into a different set of axes. The total element count must stay identical, and no value changes, only where it sits.

Broadcasting

Lets a smaller tensor be reused across a larger one during an operation, following fixed rules, so mismatched shapes can still combine without manually copying data to make them match.

The Nook of Wonder Theorems & Beautiful Patterns — the arithmetic of reshaping

A tensor's total element count is the product of every number in its shape. A reshape is only valid if that product stays the same:

shape (4, 6) → 4 × 6 = 24 elements
shape (2, 3, 4) → 2 × 3 × 4 = 24 elements

Both shapes hold 24 numbers, so reshaping one directly into the other is valid; the numbers just get relabeled across a different set of axes. A reshape into (2, 3, 5) would fail, since 2 × 3 × 5 = 30 doesn't match the original 24.

Broadcasting a shape-(6,) bias across a shape-(4, 6) matrix, one rule: axes are compared from the right, and any axis of size 1 (or missing) is treated as if it repeated to match:

matrix: (4, 6)
bias: (6,) → treated as (1, 6)
result: (4, 6), bias applied to every row

Where tensors show up

Once you know to look for shapes, they’re everywhere in how AI systems are described, even when the word “tensor” never gets said out loud. The output of an attention layer inside a transformer is a tensor. A batch of audio clips turned into spectrograms for a speech model is a tensor. The weights connecting one layer of a neural network to the next, the actual learned parameters from the very first article in this series, are stored as tensors too, usually rank 2 or higher. Nearly every number that moves through a model during training or inference is sitting inside one.

This is also a good moment to notice how much of this series has quietly been building toward the same picture from different directions. A weight is a tensor. An embedding is a vector, which is a rank-1 tensor. A dimension is one axis of that vector. None of these are separate ideas competing for attention; they’re the same underlying object, numbers organized in space, described from three different angles depending on which property is relevant at the time.

  • A batch of photos being processed together by an image recognition model
  • A stack of audio waveforms converted to spectrograms for a speech app
  • Video being fed frame by frame, in batches, into a video generation or captioning model
  • The weight matrices you learned about in Part 1, now revealed to be tensors all along
  • Not a concept limited to physics; the AI usage borrows the name but not the general-relativity baggage