Every series I’ve written so far has bumped into quantization from the side — QLoRA compressing frozen bases, the optimization guide’s memory math, GGUF files on Hugging Face with cryptic suffixes. Time to attack it head-on, because quantization is the single highest-leverage optimization in practical ML: hours of work, 2–4× less memory, often 2× faster generation, usually imperceptible quality cost. But to use it well — and to understand why it sometimes fails badly — you need to know what’s actually happening to the numbers. This first article builds quantization from the ground up.
What a weight actually is, and why 16 bits might be too many
A model’s weights are stored as floating-point numbers — most commonly BF16: 1 sign bit, 8 exponent bits (range), 7 mantissa bits (precision). Sixteen bits per weight, two bytes, times 7 billion weights: 14 GB. The bet of quantization is that most of those bits are wasted — that the information each weight actually contributes can survive in 8, 4, even fewer bits.
Why would that be true? Three converging reasons. Neural networks are trained with noisy gradients on noisy data and are robust to perturbation by construction — a weight that only works if its 7th decimal is exact would never have survived training. Second, redundancy: billions of weights encode overlapping features, so per-weight errors average out across the massive dot products of inference (a matmul summing thousands of small errors with random signs cancels most of them). Third — and this is the systems reason quantization pays speed, not just memory — LLM generation is memory-bandwidth-bound: producing each token requires streaming essentially all the weights from GPU memory through the compute units, and the arithmetic finishes long before the loading does. Halve the bytes per weight and you nearly halve the wall-clock per token, even if you do the math at full precision after loading. Memory is the tax; quantization is the tax cut.
The core mechanism: a grid and a scale
Quantization maps continuous floats onto a small integer grid. The minimal version, for 4-bit symmetric quantization of a group of weights: find the largest absolute value in the group; divide by 7 (the largest 4-bit signed integer) to get a scale; then each weight is stored as round(weight / scale) — an integer from −8 to 7 — and reconstructed at use-time as integer × scale. That’s the entire trick: store cheap integers plus one shared scale; multiply back on the way in. The reconstruction isn’t exact — every weight lands on the nearest grid point, and the rounding gap is the quantization error. An asymmetric variant adds a zero-point offset so the grid can shift to fit lopsided distributions (crucial for activations like post-ReLU values, which are all-positive); weights are usually symmetric-friendly.
Everything interesting in quantization is about managing that rounding error, and the first lever is granularity: how many numbers share one scale. One scale per tensor is cheapest and worst — a single outlier weight stretches the grid so far that ordinary weights collapse onto a few points near zero. Per-channel scales (one per output row) are much better. Modern 4-bit schemes go finer still: per-group scales, one per block of 32–128 consecutive weights, so each little neighborhood gets a grid fitted to its own range. The cost is storing the scales themselves — the metadata overhead that turns “4-bit” into an effective 4.25–5 bits per weight, and which QLoRA’s double-quantization trick (compressing the scales too) attacks. There’s a fundamental trade here worth naming: finer groups, lower error, more overhead; every real format is a point on that curve.
Second lever: the grid itself doesn’t have to be uniform. Weights are bell-curved — dense near zero, thin tails — so evenly-spaced levels waste representation where no weights live. Formats like NF4 (from QLoRA) space their 16 levels by quantiles of a Gaussian, matching the grid to the distribution. The general principle: bits are a budget; spend them where the data is.
Weights, activations, and the KV cache: three different problems
“Quantize the model” is actually three separate decisions, with very different difficulty:
Weight-only quantization (the W4A16 pattern: 4-bit weights, 16-bit activations) is the workhorse. Weights are static — quantized once, offline, with as much care as you like — and dequantized to 16-bit on the fly for each matmul. You get the full memory and bandwidth win; the compute still runs in fast, safe 16-bit. Nearly everything you download as “4-bit” (GPTQ, AWQ, GGUF files) is this.
Activation quantization (W8A8 and beyond) also quantizes the tensors flowing through the model, unlocking genuinely faster integer matmul hardware (INT8 tensor cores). But activations are computed live — scales must be chosen either from calibration statistics ahead of time (static) or on the fly (dynamic) — and, critically, activations in large LLMs contain systematic outliers that make them brutally hard to quantize. That outlier story is deep enough that it’s the entire subject of my next article.
KV-cache quantization targets the third memory consumer: the attention cache, which for long contexts and big batches can rival the weights. 8-bit KV is close to free quality-wise and is often the difference between fitting a workload and not; 4-bit KV is usable with per-head care. If you serve long-context anything, this knob matters as much as the weight format.
Calibration, and how quality is measured
Beyond round-to-nearest, quality comes from choosing scales (and roundings) intelligently — and for that you need to know which errors matter. Enter the calibration set: a few hundred representative text samples run through the model to record activation statistics. Even basic PTQ (post-training quantization) uses calibration to clip ranges sensibly — because protecting one extreme outlier weight at the cost of precision on ten thousand ordinary ones is a bad trade, and slightly clipping the range (accepting large error on the rare extremes for a finer grid everywhere else) is often optimal. The sophisticated methods in the next article use calibration far more aggressively — but the humble version of the lesson is already here: quantization error only matters where it changes the output, and calibration is how you find out where that is.
Which raises the measurement question I want to flag now and return to in article 3, because it’s where practitioners get burned: the standard quick metric is perplexity on held-out text, and it’s genuinely useful for ranking formats — but small perplexity gaps can hide meaningful damage on specific capabilities, with the recurring finding that multi-step reasoning, math, and code degrade first while fluent surface text degrades last. A quantized model that reads perfectly and quietly fails at arithmetic is the characteristic failure. Perplexity for triage; task evals — your tasks — for decisions.
Where this leaves us
The picture so far: quantization works because networks are robust and redundant, it pays because inference is bandwidth-bound, and its craft lives in three nested choices — grid shape, scale granularity, and which tensors to touch. With just this toolkit (round-to-nearest, per-group scales, calibration-based clipping), 8-bit weights are essentially free and naive 4-bit is… almost good. What stands between “almost” and the excellent 4-bit models everyone actually uses is a strange empirical discovery about large transformers — a handful of hidden dimensions with values a hundred times larger than everything else, which shatter naive quantization and forced the field to invent cleverer machinery. That discovery, and the family of methods built around it — LLM.int8(), SmoothQuant, GPTQ, AWQ — is the next article, and it’s the best detective story in the efficiency literature.