Shrinking a 2.8-Trillion -Parameter Model
Kimi K3 shipped as the largest open-weight model anyone has released. This piece looks at what the compression crowd has actually managed to do about that, two days in.
Read entryFig. I — Filtered by tag
8 meditations tagged quantization. Back toall tags or the full archive.
Kimi K3 shipped as the largest open-weight model anyone has released. This piece looks at what the compression crowd has actually managed to do about that, two days in.
Read entryDeep Dives
Moonshot AI just shipped the biggest open-weight model anyone's ever released, and it's not just big for the sake of being big — there's real engineering under the hood.
Read entryTraining a capable model is only half the battle. The model that comes out of pretraining or fine-tuning is almost never the model you actually want to ship.
Read entryQuantization · Part 4
Everything in this series so far has been post-training quantization: take a finished model, compress it, hope the damage is small.
Read entryQuantization · Part 3
Open any popular model's page on Hugging Face and you'll find the GGUF listings: Q4KM, Q5KS, Q6K, Q80, IQ2XS — a wall of cryptic suffixes, each a different point on a size-quality curve, downloaded millions of times by people running models on gaming PCs an…
Read entryQuantization · Part 2
Here's a puzzle that stumped the field around 2022. Quantization recipes that worked beautifully on small language models — clean INT8, minimal quality loss — fell off a cliff on big ones.
Read entryQuantization · Part 1
Every series I've written so far has bumped into quantization from the side — QLoRA compressing frozen bases, the optimization guide's memory math, GGUF files on Hugging Face with cryptic suffixes.
Read entryLoRA Deep Dive · Part 2
LoRA solved half the memory problem: with adapters, the trainable state — gradients and optimizer moments — collapses to almost nothing.
Read entry