Here’s a puzzle that stumped the field around 2022. Quantization recipes that worked beautifully on small language models — clean INT8, minimal quality loss — fell off a cliff on big ones. Somewhere around 6–7B parameters, the same techniques that were nearly lossless at 1B suddenly produced gibbering wrecks. It wasn’t gradual degradation; it was a phase transition. The investigation of that cliff produced one of my favorite empirical discoveries in the literature, and every major quantization method since — LLM.int8(), SmoothQuant, GPTQ, AWQ — is best understood as a different answer to what the investigation found. This article is that detective story.
The discovery: emergent outlier features
Tim Dettmers and colleagues traced the failure to something strange in the activations of large transformers: a tiny number of hidden dimensions — a handful out of thousands — carrying values up to 100× larger than everything else. And these weren’t random spikes: the same few dimensions fired huge values consistently, across tokens, across layers, in every sufficiently large model they examined. Below ~6B parameters the phenomenon was mild and scattered; past it, the outlier dimensions became systematic and universal. An emergent property of scale.
Why this murders quantization follows directly from article 1: a quantization grid must span the range of the values it covers. One dimension hitting ±100 while its neighbors live in ±1 forces the shared scale to stretch across ±100 — leaving the ordinary values, which carry most of the information, crushed onto two or three grid points around zero. The outliers survive; everything else is rounded into mush. Worse, ablations showed the outlier dimensions are disproportionately important — zero them out and model quality craters — so you can’t clip them away either. Big models thus present quantization with a trap: the values you must preserve precisely are the ones that destroy precision for everything else. Every method below is an escape from this trap, and they’re worth learning as a set because they escape in four genuinely different directions.
Escape #1 — Quarantine: LLM.int8()
Dettmers’ own answer is the most direct: separate the two populations. At runtime, identify the outlier dimensions (a fraction of a percent of columns), route those through the matmul in full 16-bit precision, and quantize everything else to INT8 with per-vector scales. Two matmuls — one huge and cheap, one tiny and precise — summed at the end. Quality: essentially lossless, at every scale tested, with no calibration or tuning. The costs: mixed-precision decomposition adds overhead (early implementations were often slower than FP16 — memory saved, speed lost), and it’s an 8-bit method, not a path to 4. Its historical importance is huge — it made big models runnable on small GPUs and diagnosed the disease — but its deeper legacy is the principle: treat outliers as a separate species, not as noise.
Escape #2 — Relocation: SmoothQuant
SmoothQuant starts from a sharper observation: the outlier problem lives in activations; the weights of those same channels are perfectly tame. And there’s a free mathematical move available — for any per-channel scaling s, (X/s)·(s·W) = X·W. So: divide the activation outlier channels down by a smoothing factor and multiply the corresponding weight channels up by the same factor. Nothing changes numerically; the difficulty migrates from activations (hard to quantize, computed live) into weights (easy to quantize, prepared offline). After smoothing, both sides are well-behaved enough for straight W8A8 — full INT8 matmuls on integer tensor cores, which is where the real serving speedups live. The elegance here is what I love: no extra compute, no mixed precision — just relocating variance to where it’s cheap to handle. The α knob (how much difficulty to migrate) is the method’s one piece of tuning, and its limit is honest: it redistributes the outlier problem rather than eliminating it, which is why W8A8 works and W4A4 still doesn’t.
Escape #3 — Compensation: GPTQ
GPTQ attacks a different formulation: forget activations — get weights down to 4 or 3 bits with minimal damage. Its ancestor (Optimal Brain Surgeon lineage) contributes the key idea: when you round one weight, you inflict a known error — and the remaining unquantized weights can be adjusted to absorb it. GPTQ does this at scale: layer by layer, quantize weight columns one at a time, and after each rounding, update all not-yet-quantized weights (using second-order information — a Hessian built from calibration activations) to compensate for the error just made. It’s error-feedback, like dithering in audio: each individual rounding is a small lie, but the running compensation keeps the layer’s overall output faithful. Clever engineering (lazy batched updates, Cholesky tricks) made this tractable for billion-parameter models in hours on one GPU. GPTQ is what made 3–4 bit weights of huge models actually good, and its machinery — calibration-driven, layer-wise output matching — became the template for a generation of methods. Its dependence on calibration data is also its caveat: quantize against wikitext, and quality can wobble on distributions far from it.
Escape #4 — Protection: AWQ
AWQ’s founding observation refines the outlier story one more step: what matters isn’t which weights are large — it’s which weights meet large activations. Salience is a property of the product. Their striking demo: keep just ~1% of weight channels — chosen by activation magnitude — in FP16, and 4-bit quality jumps dramatically; choose that 1% by weight magnitude and it barely helps. But mixed precision is hardware-unfriendly, so AWQ converts protection into scaling (the same trick as SmoothQuant, aimed at a different goal): multiply salient weight channels up before quantization (dividing activations correspondingly), so the important channels occupy more of the quantization grid and suffer proportionally less rounding error. No backprop, light calibration (it uses only average activation scales, making it less prone to overfitting the calibration set than reconstruction-based methods), fast to produce, and — with its optimized kernels — fast to run. That balance of quality, robustness, and speed is why AWQ became arguably the default 4-bit format for GPU serving.
The pattern, and the scorecard
Line them up and the intellectual structure is beautiful — four verbs for one noun: quarantine the outliers (LLM.int8()), relocate them (SmoothQuant), compensate for rounding damage (GPTQ), protect what matters most (AWQ). All four rest on the same discovery: quantization error is not uniform in importance — a tiny structured minority of values carries disproportionate weight, and knowing where (via calibration) is worth more than any amount of cleverness about the grid.
The practical scorecard I’ve distilled: serving on GPUs at 4-bit → AWQ or GPTQ, with AWQ the robustness-favoring default and GPTQ often a hair better when carefully calibrated in-domain; need integer-matmul throughput at scale → SmoothQuant-style W8A8 (now standard inside serving stacks like TensorRT-LLM); want zero-risk 8-bit with no calibration → the LLM.int8() lineage. And hovering over all of it, one more contender I’ve deliberately postponed: the GGUF k-quant family that powers llama.cpp and most local AI — a different philosophy (no reliance on calibration, hierarchical block structure, CPU-first) that deserves its own article, because it’s the format most people actually touch. That’s next.