Quantization · Part 4

Below Four Bits: QAT, BitNet, and the Race Toward One-Bit Intelligence

Everything in this series so far has been post-training quantization: take a finished model, compress it, hope the damage is small.

Good Omens Studio7 min readMachine Learning

Everything in this series so far has been post-training quantization: take a finished model, compress it, hope the damage is small. That paradigm has a floor. Around 3–4 bits, PTQ methods — however clever their outlier handling — start visibly losing the model, and by 2 bits even the best of them are performing triage. Yet models at 2 bits, 1.58 bits, even 1 bit exist and work. The trick is a change of philosophy: stop compressing models after training and start training models that are born quantized. This final article of the series (and of the whole project) is about that frontier — quantization-aware training, the BitNet result that startled everyone, and what an ultra-low-bit future implies about hardware and the field.

Why PTQ hits a floor

The intuition, assembled from the previous three articles: PTQ works because trained networks are robust to small perturbations, and 8-bit or 4-bit rounding is a small perturbation — errors average out across big dot products, and calibration-guided methods steer the residual damage away from what matters. But precision loss compounds exponentially as bits drop: each bit removed doubles the coarseness of the grid. At 4 bits, 16 levels can still trace the shape of a weight distribution; at 2 bits, four levels cannot — the “perturbation” is no longer small relative to the signal, entire distinctions between weights collapse, and no amount of post-hoc cleverness can recover information the grid simply cannot express. The model you trained assumed a continuum; you’re forcing it onto a lattice it never agreed to.

The fix follows from stating the problem that way: let the model agree to the lattice during training.

QAT: training with the constraint switched on

Quantization-aware training simulates quantization inside the training loop: on each forward pass, weights (and possibly activations) are “fake-quantized” — rounded to the target grid — so the loss reflects the quantized model’s true behavior. The obstacle is that rounding has zero gradient almost everywhere; backprop through a staircase learns nothing. The standard workaround, crude and effective, is the straight-through estimator (STE): use the quantized values in the forward pass, but pretend the rounding was the identity function in the backward pass, letting gradients flow to an underlying full-precision “shadow” copy of the weights. The shadow weights accumulate small updates in high precision; the quantized projection of them is what the network actually experiences. Over training, the model migrates toward weight configurations that work well on the grid — sitting in wide, flat minima where rounding doesn’t matter, routing information away from precision-hungry pathways.

The economics: QAT costs real training compute (though “finish with a QAT phase” — a fraction of full training — captures most of the benefit, and QAT-finished releases of open models are increasingly common). What it buys is exactly the sub-4-bit regime: at 8 bits QAT is unnecessary, at 4 bits it’s a modest improvement over the best PTQ, and at 2–3 bits it’s the difference between degraded and genuinely good. The lower the target, the more training has to know.

BitNet: the 1.58-bit provocation

Then push QAT to its logical extreme. Microsoft Research’s BitNet b1.58 (2024) trains transformers whose weights take exactly three values: −1, 0, +1 — log₂(3) ≈ 1.58 bits of information per weight — from scratch, quantized from the first step (with higher-precision activations, ~8-bit, and full-precision shadow weights during training only). The startling claim, which follow-up replications and scaled-up releases have substantially supported: at moderate scale (a few billion parameters), ternary models match full-precision transformers of equal size on language modeling and downstream tasks. Not “close enough” — matching, within noise, while using ~10× less weight memory and dramatically less energy.

Two things about this deserve real reflection.

First, the arithmetic consequence: multiplying by −1, 0, or +1 isn’t multiplication — it’s sign-flipping and skipping. Matrix multiplication, the operation that consumes the overwhelming majority of AI compute, reduces to addition. Multipliers are the expensive, power-hungry part of arithmetic silicon; an architecture that eliminates them isn’t an optimization of current hardware — it’s an argument for different hardware: adder-centric accelerators radically cheaper and cooler than multiply-based GPUs. (On today’s chips, ternary models already run impressively via bit-packing tricks — official CPU inference frameworks demonstrate large models at readable speeds on laptops — but the deep payoff is co-designed silicon, and that flywheel turns slowly.)

Second, the scientific consequence, which loops back to the very first thing this whole blog series established: networks are overparameterized, and the information a model actually uses is far less than its container. BitNet quantifies the claim shockingly tightly — ~1.58 bits per weight suffices, if the model learns within that constraint from the start. Capability, it increasingly seems, lives in the structure — which connections exist, their signs, the patterns across billions of them — far more than in the precision of any individual value. Precision, in this view, was mostly scaffolding for optimization: gradients need a smooth continuum to descend; the destination doesn’t need it. Which is also the honest statement of BitNet’s limits: you can’t convert an existing model to ternary (the knowledge is written in the precision you’d be deleting — PTQ to 1.58 bits fails hard); it demands training from scratch, which is why the world’s enormous investment in existing full-precision checkpoints keeps 4-bit PTQ dominant in practice while ternary matures. And the strongest evidence remains concentrated at small-to-mid scale; whether the parity claim holds cleanly at frontier scale is still the open bet.

Between 4 and 1.58: the PTQ frontier fights back

The extreme-low-bit space isn’t only QAT’s. The codebook school — QuIP#, AQLM, and kin — pushes post-training quantization to 2–3 bits by abandoning independent per-weight rounding entirely: quantize weights in groups against learned or lattice-structured codebooks (after “incoherence” rotations that scramble outliers into well-behaved distributions — the outlier problem solved by change of basis, a fourth escape to add to article 2’s taxonomy). These methods, plus hybrid recipes that fine-tune briefly after aggressive quantization, keep narrowing the gap — 2-bit versions of large models that are convincingly usable. The pattern of the whole field in one sentence: every bit removed demands that more global structure be taken into account — from single weights, to blocks, to channels-with-calibration, to codebooks-with-rotations, to the entire training trajectory.

Closing the series — and the project

This is the seventeenth and last article of the project I set out on: four series deep into how models learn, how they’re aligned, what they’re fed, how they’re specialized — and now, how far they compress. If I compress that (fittingly) into one thought, it’s this: at every layer of the stack, the field keeps rediscovering the same asymmetry — what a model needs to be is much smaller than what it needs to become. Training demands abundance: parameters, precision, data, exploration. The capable artifact at the end demands astonishingly little: a low-rank steer selects its behavior, a thousand curated examples set its style, four bits carry its knowledge, three values per weight might suffice for its structure. Nearly everything we spend is scaffolding for the search; nearly everything we keep is the small, sharp thing the search found.

The big optimization guide I posted at the start of this journey covers how these pieces — distillation, quantization, pruning, adapters, serving — compose in practice. Having now dug the foundations under each of them, I read my own guide differently: less as a bag of tricks, more as one repeated move. Find what’s essential. Throw away the rest. It’s not a bad philosophy off the clock, either.