The Bit Explainers

The Chip That Learned to Think in Parallel

Why the hardware running today's AI was built to draw video game triangles, and what happens when you point a few thousand of them at the same problem.

Good Omens Studio18 min readMachine Learning

Every explainer in this series has talked about models getting bigger, training on more data, running faster at inference — without stopping to ask what’s actually doing the arithmetic underneath all of it. For nearly every serious model built since 2012, the answer is a graphics card. Not a chip designed for reasoning or language or anything resembling thought — a chip designed to paint pixels on a screen sixty times a second, repurposed almost by accident into the substrate the entire field now runs on. Understanding why that repurposing worked, and where it’s headed next, tells you more about the shape of modern AI than most single concept in this series so far.

None of this was planned in advance. The hardware that trains today’s largest models descends from chips built to sell arcade cabinets and home consoles, refined for two decades by an industry that cared about frame rates and nothing else. The fact that this lineage turned out to be exactly what AI needed is either an enormous coincidence or a sign that “many small, identical calculations, done independently” is a far more common shape of problem than it first appears. This article traces that lineage from arcade hardware to the data centers being built today, and follows the roadmap forward to see where the same design logic is taking the industry next.

Built to paint, not to think

Rewind further than most GPU histories bother to: the earliest dedicated graphics chips showed up in 1970s arcade cabinets, doing nothing more sophisticated than moving a handful of sprites around a screen. By the mid-1990s, consumer 3D accelerator cards — 3dfx’s Voodoo line is the one most often credited with kicking off the category — arrived to handle a much harder version of the same job: rendering full three-dimensional scenes, in real time, for PC games. That’s the lineage. Nothing about it was designed with scientific computing, let alone neural networks, in mind. It was designed to sell more copies of Quake.

Video games wanted 3D worlds, and 3D worlds are made of triangles — millions of them, redrawn from scratch every fraction of a second as the camera moves. Turning a triangle into the colored pixels you see on screen involves a chain of fairly simple steps: figure out where its corners land after the camera moves, work out which pixels fall inside it, then compute a color for each of those pixels based on lighting, texture, and material. None of that is intellectually hard math. It’s the same handful of multiplications and additions, over and over, for every triangle and every pixel, every single frame.

The part worth sitting with is that each of those pixel calculations doesn’t need to know what any other pixel is doing. Pixel (400, 220) and pixel (401, 220) both run the identical shading formula, but neither one waits on the other’s answer. That’s a very unusual property for a computing problem to have, and it’s the property that decided what a GPU would look like as a piece of silicon.

A processor built to run one instruction stream as fast as possible — fetch, decode, predict what comes next, execute, repeat — is optimized for a different kind of problem: one long chain of steps where each step depends on the last. Rendering pixels is the opposite. It’s an enormous pile of short, identical, independent jobs. So graphics chip designers made a bet that would later turn out to be one of the more consequential engineering decisions of the last thirty years: instead of a few large, clever cores, build thousands of small, deliberately simple ones, and hand each of them one pixel’s worth of arithmetic at a time.

Two workers, two philosophies

That design choice is easiest to see next to the chip everyone already has some intuition for. A CPU and a GPU aren’t a fast version and a slow version of the same thing — they’re built around opposite bets about what kind of work is coming.

CPU — few, powerful workers

A handful of cores (commonly single digits to a few dozen), each packed with branch prediction, out-of-order execution, and deep caches. Built to run one complicated, unpredictable sequence of instructions as fast as possible — the logic in an operating system, a database query, a web server deciding what to do next.

GPU — thousands of simple workers

Thousands of small cores, each stripped of most of that overhead, built to run the same instruction on many pieces of data at once. Weak at anything that requires branching or guesswork about what comes next, extremely strong at the same arithmetic op repeated at massive scale.

Neither design replaces the other, which is why every AI training run still needs a CPU nearby — something has to load the data, manage the program, and hand off the actual number-crunching to the GPU sitting next to it. The CPU is the conductor; the GPU is several thousand musicians who only know how to play one note, in perfect time, on command.

The gap shows up in the numbers, not just the description. A high-end consumer CPU might carry somewhere around 8 to 24 cores, each with its own large chunk of fast cache memory and the ability to run a completely different program from its neighbor. A single high-end GPU carries thousands of cores — grouped into clusters that all execute the same instruction together, sharing a much smaller pool of cache per core, because the whole point is that they don’t need to remember much individually. A CPU core is a generalist with a good memory. A GPU core is a specialist with almost none, relying on there being thousands of identical specialists nearby doing the same job on the neighboring piece of data.

The accident that changed everything

For most of the 2000s, GPUs were still understood purely as graphics hardware. But a handful of researchers doing scientific computing noticed something odd: a lot of physics simulation and linear algebra involved the exact same shape of problem as shading pixels — huge grids of numbers, each one updated by an identical, independent calculation. Some of them started disguising their math as graphics, encoding matrices as textures and equations as shading routines, just to get access to that parallel hardware. It worked, but it was a hack, and an unpleasant one to program.

NVIDIA’s response was CUDA, released in 2006: a way to write general-purpose parallel programs for a GPU directly, without pretending they were graphics at all. At the time this looked like a modest developer-tools decision, and for years it mostly was — CUDA spent its first half-decade as a niche tool for a small community of scientific computing researchers, with no obvious mass market and no guarantee it would ever be more than that. NVIDIA kept investing in it anyway, through several product cycles where the payoff wasn’t visible yet. In hindsight it was the moment the GPU stopped being exclusively a graphics chip and became a general parallel computer that also happened to be good at graphics, but that wasn’t obvious to anyone at the time, including NVIDIA.

The moment that made the bet pay off arrived in 2012. Alex Krizhevsky, Ilya Sutskever, and Geoffrey Hinton trained a convolutional neural network — later known as AlexNet — on two consumer NVIDIA GPUs, and entered it into the ImageNet image recognition competition. It didn’t just win; it beat the next-best entry by a margin large enough that the rest of the field took notice immediately. The reason it worked wasn’t a smarter algorithm alone — convolutional networks had existed for years. It was that training a neural network is, underneath all the terminology, dominated by the exact same operation graphics chips were already built to devour: multiplying and summing huge grids of numbers, independently, over and over.

That single result reoriented an entire industry’s roadmap. Every major deep learning framework built afterward — TensorFlow, PyTorch, and everything since — was written to run on CUDA first, and everything else second. That’s not a small detail. It means the six years NVIDIA spent building developer tools for a niche audience turned into the deepest software moat in the hardware industry: switching away from NVIDIA doesn’t just mean buying different chips, it means rewriting or recompiling software written against an ecosystem competitors have spent nearly two decades trying, without full success, to match.

Learning needs multiplication, a lot of it

To see why that overlap is so exact, it’s worth pulling apart what a neural network is actually doing when it trains, and why that operation happens to be the one GPUs were already shaped for.

In one sentence

Training a neural network is mostly the same operation — multiplying and adding two grids of numbers together — repeated billions of times, and that operation splits cleanly across thousands of independent workers.

Intuition

Picture a spreadsheet with millions of cells, where each cell's value is found by multiplying one row of numbers against one column of numbers and adding up the results. Every cell's answer depends only on its own row and column — never on any other cell's answer at the same time. So instead of one accountant filling the spreadsheet in one cell at a time, you can hand every cell to its own accountant and let all of them work simultaneously, then just collect the finished sheet.

Technical

A neural network layer's forward pass, and most of its backward pass during training, reduces to matrix multiplication: multiplying an input matrix by a weight matrix to produce an output matrix. Each entry of the output is a dot product of one row and one column from the inputs, computed with no shared state between entries. That independence is exactly what SIMT execution — the model GPU cores run on, where many cores execute the same instruction on different data simultaneously — was built to exploit. Backpropagation, the process that computes how much to adjust each weight after a mistake, is built from the same operation run in reverse: it multiplies the same matrices, transposed, against the error signal flowing backward through the network. Learning and predicting turn out to lean on the identical piece of arithmetic, which is part of why one kind of chip could serve both jobs.

So why not just build a CPU with thousands of cores instead of buying a separate kind of chip?

Because CPU cores carry a lot of machinery a matrix multiply doesn't need: branch prediction, speculative execution, deep instruction pipelines, all built to guess correctly about unpredictable, sequential code. That machinery costs transistors and power per core. A GPU core strips almost all of it out and spends the freed-up space on more cores instead — trading the ability to run arbitrary, branchy logic well for the ability to run the same simple instruction across an enormous number of data points at once. For a neural network, which is arbitrary logic almost never and matrix multiplication almost always, that trade is close to free.

The Nook of Wonder Theorems & Beautiful Patterns — counting the multiplications in a single matrix multiply

For a matrix A of size m × k multiplied by a matrix B of size k × n, producing result C of size m × n:

FLOPs ≈ 2 · m · n · k

Each of the m × n output entries needs k multiplications and k − 1 additions — roughly 2k operations per entry, hence the factor of 2.

4096 × 4096 layer: 2 · 4096 · 4096 · 4096 ≈ 137 billion FLOPs

That's one matrix multiply, inside one layer, for one forward pass — a size roughly typical of a hidden layer in a mid-sized transformer. A single training step touches this many layers, forward and backward, for every example in the batch. The count is why "how many FLOPs did this model train with" became a headline number in its own right — it's just this formula, applied billions of times over.

Built for tensors: the specialization era

Once it was clear that AI training was going to be the dominant use of this hardware, GPU makers stopped treating matrix multiplication as something the chip happened to be okay at, and started designing directly for it. Three changes did most of the work.

The first was dedicated circuits for exactly this operation. Starting with NVIDIA’s Volta architecture in 2017, GPUs began shipping Tensor Cores — units that compute a small matrix multiply-and-add in a single hardware step, instead of stepping through it as a sequence of ordinary instructions the way a general-purpose core would. It’s the difference between a calculator with a dedicated “multiply-accumulate” button and one where you punch in multiply, then equals, then plus, then equals, by hand — same result, far fewer steps, and steps are what cost time.

The second was giving up precision the model doesn’t need. Early scientific computing insisted on 32-bit or 64-bit numbers for accuracy. Neural networks, it turns out, tolerate much coarser numbers without losing much — 16-bit, 8-bit, and increasingly 4-bit formats are now standard for large parts of training and inference. Fewer bits per number means more numbers fit through the same wire and the same core in the same amount of time, so this single trade-off has been worth roughly as much raw speed over the last decade as the architecture changes themselves.

The third was recognizing that cores sitting idle waiting for numbers to arrive aren’t actually fast, no matter how many of them there are. That shifted a huge amount of engineering effort toward memory: stacking memory chips directly on top of or beside the compute die (a design called HBM, high-bandwidth memory) to shrink the distance data has to travel, and building extremely fast links — NVIDIA calls its version NVLink — so that dozens or hundreds of GPUs can share memory and work almost as if they were one much larger chip.

A quieter fourth trend runs alongside those three: exploiting sparsity, the fact that a large share of the numbers inside a trained network end up close to zero and contribute almost nothing to the final answer. Recent Tensor Core designs can detect a structured pattern of zeros in a matrix and simply skip multiplying them, which is free performance as long as the skipping logic itself doesn’t cost more than the multiplication it avoids. Combined with lower-precision formats — FP16 and BF16 through the 2010s, FP8 more recently, and FP4 now standard on the newest chips — the net effect is that a huge amount of the last decade’s GPU speedup has come from doing less arithmetic per number and skipping arithmetic that doesn’t matter, not from making each multiplication happen faster.

Per-GPU memory bandwidth by generation — how fast data can reach the cores, not how fast the cores can compute. Figures per NVIDIA’s published architecture specs.

The big plan: from chip to planet-scale computer

The unit of progress in this field has quietly changed. A few years ago, the interesting question was how fast a single GPU was. Today, NVIDIA and its competitors design, ship, and sell entire racks as if they were one product — dozens to over a hundred GPUs, wired together with NVLink so tightly that software can treat them as a single enormous accelerator rather than a cluster of separate machines. NVIDIA’s current generation of these rack-scale systems, Vera Rubin, pairs a new “Rubin” GPU with a companion “Vera” CPU and is shipping to cloud providers in the second half of 2026, with a further “Rubin Ultra” generation already confirmed for 2027 and a “Feynman” generation penciled in for 2028. That’s a new architecture roughly every year — a pace of hardware iteration that has no real precedent outside this industry.

Architecture Ships Dense FP4 compute (per GPU) Memory Memory bandwidth
Blackwell B300 2025 15 petaFLOPS 288 GB HBM3E 8 TB/s
Vera Rubin (R200) H2 2026 50 petaFLOPS 288 GB HBM4 13 TB/s
Rubin Ultra H2 2027 (planned) ~100 petaFLOPS 1 TB HBM4e ~32 TB/s

Two forces are pushing that cadence. One is demand: the amount of compute used to train frontier models has been growing far faster than any single chip generation could keep up with on its own, so scaling now happens by adding more chips, wired closer together, as much as by making each chip individually better. The other is a bottleneck almost nobody outside the industry was tracking a decade ago — electricity. A single rack of Vera Rubin-class hardware draws hundreds of kilowatts, and data centers are increasingly described by planners in gigawatts rather than server counts. The “big plan,” as NVIDIA’s own roadmap frames it, isn’t really about a faster chip anymore. It’s about treating the chip, the network connecting it to its neighbors, the cooling system, and the power feeding all of it as one designed system, refreshed on an annual clock.

NVIDIA isn’t alone in this race, even though it’s the name that dominates headlines. AMD ships a competing line of accelerators under its Instinct brand. Google has spent over a decade building its own AI chips, called TPUs, for internal use and cloud customers, taking a related but distinct architectural approach. Amazon and Microsoft have both started designing custom AI silicon for their own data centers, partly to reduce dependence on any single supplier. None of these has displaced NVIDIA’s GPU-and-CUDA ecosystem as the default, but the fact that trillion-dollar companies keep building alternatives anyway says something about how central this piece of hardware has become to everything built on top of it.

The power problem is worth taking seriously rather than treating as a footnote, because it’s reshaping decisions that used to be purely about chip design. Packing thousands of Tensor Cores and stacks of high-bandwidth memory into a small area generates enormous amounts of heat, and the industry has largely run out of room to solve that with fans. The newest racks are liquid-cooled by default, with coolant piped directly to cold plates sitting on top of the chips, because moving that much heat through air alone stopped being physically practical a couple of GPU generations ago. Data centers built for this hardware are now sited and negotiated for around electricity access the way factories once were sited around rivers, and it’s becoming common for a single new AI data center campus to require its own dedicated power generation rather than drawing off the existing grid. None of that is a side story to the compute story. It’s the same story: a chip built to do one narrow kind of arithmetic extremely well is only as useful as the electricity, cooling, and networking someone builds around it, and increasingly the biggest engineering constraint on how much AI compute exists isn’t the chip design at all.

Where you've met this already

The overlap between “graphics chip” and “AI chip” shows up in more everyday places than it might seem.

  • The portrait-mode blur and computational photography on your phone run on the same small GPU that renders its games — mobile chips carry a scaled-down version of this same parallel design.
  • Video editing software renders effects and exports footage faster with a GPU because color correction and filters are, again, the same operation applied independently to every pixel.
  • The 2017–2021 cryptocurrency mining boom emptied GPU shelves worldwide for the same underlying reason AI training does: mining computes the same hash function over and over, independently, at massive scale.
  • Every image, video, or text response a generative AI tool produces for you is, at the hardware level, a very long sequence of the matrix multiplications this article walked through.
  • A GPU is not simply "a faster CPU" — swap that assumption in and the rest of this article, and most of how modern AI hardware is designed, stops making sense.