The Bit Explainers · Deep Dive · Follow-Up

Shrinking a 2.8-Trillion -Parameter Model

Kimi K3 shipped as the largest open-weight model anyone has released. This piece looks at what the compression crowd has actually managed to do about that, two days in.

Good Omens Studio10 min readMachine Learning

Last time, we covered what Kimi K3 actually is: 2.8 trillion total parameters, a mixture-of-experts model that only wakes up 16 of its 896 experts per token, and open weights that Moonshot AI put up for public download on July 27. The number that gets repeated is 2.8T. The number that matters for actually running it is closer to 50B active parameters per token, though even that smaller slice still has to sit in memory next to the rest of the model, waiting to be picked. That’s the question this piece is chasing: what is anyone actually doing to make Kimi K3 smaller?

The honest answer is that “smaller” is doing three unrelated jobs at once, and mixing them up is where most of the confusion online is coming from.

Already Spent

Quantization Already Happened

With Kimi K2, the pattern was familiar: Moonshot released a full-precision model, and the community spent the following weeks squeezing it down, building INT4 quants and eventually 1-bit GGUF versions that traded some accuracy for a lot less disk space. That pipeline is why a 1-trillion-parameter model could eventually run, slowly, on a single well-stocked workstation.

K3 skips that step, because Moonshot did the compression during training instead of after release. The model applies quantization-aware training from the supervised fine-tuning stage onward and ships natively as MXFP4 weights with MXFP8 activations. It arrives at 4 bits already, with no full-precision version underneath to shrink down from.

In one sentence

Quantization-aware training bakes low precision into the model while it's still learning, instead of rounding a finished model down afterward.

Intuition

Post-training quantization is like writing an essay with a fountain pen and then photocopying it at low resolution. Whatever detail the copy loses, the writer never planned for while writing. Quantization-aware training is closer to writing that same essay with a thick marker from the first draft, choosing every word already knowing exactly how much detail will survive the copy.

Technical

During training, the forward pass simulates MXFP4 rounding so gradients push the weights toward values that survive that rounding well. Post-training quantization, or PTQ, works the opposite direction: it takes a model trained entirely in BF16 and snaps its finished weights onto a low-precision grid after the fact, which is why PTQ always costs some accuracy the model was never trained to tolerate.

So if K3 is already 4-bit, can't people just push it to 2-bit or 1-bit the way they did with K2?

Not with the same payoff. K2 was trained in full precision, so it had headroom nobody had used yet, and that headroom is exactly what those 1-bit GGUF builds were spending. K3's 4-bit weights are the precision it was trained to think in, so squeezing further means rounding values that were already chosen to be as tight as they could be. There's no second discount sitting underneath the first one.

Kimi K2 (2025)

Trained and released in full precision (BF16). Post-training quantization by the community carved it down from roughly 2TB to as little as 240GB, with real quality trade-offs at the lowest bit depths.

Kimi K3 (2026)

Trained quantization-aware and released natively at MXFP4. A lossless full-precision build already runs about 1.56TB, and there's no untapped full-precision version underneath left to compress.

Cutting Experts

Pruning: Removing Whole Experts

If rounding numbers further doesn’t buy much, the next lever is removing entire pieces of the model. A mixture-of-experts model is built from many separate expert sub-networks plus a router that decides, per token, which handful of experts get to weigh in. That structure opens up a different kind of surgery: instead of shrinking every weight a little, you can remove whole experts and leave the rest of the model untouched.

In one sentence

REAP, short for Router-weighted Expert Activation Pruning, deletes the experts a model's own router rarely picks and gets little benefit from when it does.

Intuition

Picture a restaurant kitchen with ninety-six line cooks. A few of them get every third ticket, and their dishes change the meal. Most get the occasional order and turn out something forgettable. Laying off the second group shrinks payroll without changing what diners actually taste, as long as the kitchen is honest about which cooks are which.

Technical

For every expert in every MoE layer, REAP computes a saliency score from two signals: how often and how strongly the router's gate selects that expert, and how much that expert's output actually shifts the layer's result when it does fire. The lowest-scoring fraction of experts gets removed uniformly across all MoE layers, while the router's weights over the surviving experts stay untouched. That preserves the coordination between router and experts, which is exactly what earlier merging-based compression methods tended to wreck.

Doesn't deleting a third of the experts just make the model worse at everything?

Less than the raw percentage suggests, because MoE training tends to leave real redundancy on the table. Plenty of experts learn overlapping things during pretraining, so losing the least-used ones costs less than you'd guess. Cerebras has already shown this on Kimi's own sibling models: REAP cut Kimi-Linear-48B down to a 35B checkpoint at 30 percent pruning while holding accuracy close to the original on coding and tool-use benchmarks, and separate community work pruned 50 percent of experts from a quantized Kimi K2.6 checkpoint with mostly preserved short-form quality. Nobody has published a REAP checkpoint for K3 itself yet, since the model is only two days old, but the results on its own architecture family are why people expect one soon rather than doubt it's possible.

Solid bars are shipped checkpoints. The dashed K3 bar is a projection based on K2’s roughly 30% REAP cut. No pruned K3 checkpoint has been published yet.

The Nook of Wonder Theorems & Beautiful Patterns — how REAP scores an expert

S(e) = mean( gate(e) ) × ‖ output(e) ‖
prune the lowest-scoring fraction, layer by layer

gate(e) is how strongly the router weighted expert e across a calibration batch. output(e) is the magnitude of what that expert actually contributed when it was chosen. An expert the router almost never picks, or one that barely nudges the result when it is picked, ends up with a low S(e) either way, and that's the set REAP removes.

A Different Model Entirely

Distillation: Growing a Smaller Model, Not Shrinking the Big One

Quantizing and pruning both start from K3’s own weights and take something away: precision in one case, whole experts in the other. Distillation doesn’t touch K3’s weights at all. It uses K3 to teach a separate, much smaller model how to behave.

In one sentence

Distillation trains a small, separate model to imitate what a large model produces, rather than compressing the large model itself.

Intuition

An apprentice doesn't get handed the master's toolbox. They spend months watching the master work, then go build their own version with simpler tools: smaller, faster, good enough for most of the same jobs, but never a physical copy of the original.

Technical

The large model runs inference across large volumes of prompts, things like coding problems, reasoning chains, and agentic tool-use transcripts, and its outputs become training data. A dense student model, typically in the 7B to 30B range, is then fine-tuned on that data to reproduce the teacher's behavior on similar tasks. None of the teacher's weights get reused, merged, or reduced. The student is a new model that happens to have learned from a bigger one's homework.

If distillation gets you an actually small model, isn't it just better than pruning or quantizing?

It's a different trade, not a strictly better one. A distilled 14B student is genuinely small and fast, and prior work on this kind of transfer, including NVIDIA's Minitron distillations and Microsoft's work distilling from GPT-class teachers, shows students that sometimes beat their teacher on individual benchmarks while training on a small fraction of the data. But the student only learned what it was shown, so its judgment is bounded by the tasks and transcripts used in distillation. A pruned or quantized K3, however much smaller, is still K3's own judgment everywhere, including situations nobody thought to put in the distillation set.

Two Days In

Where the Plumbing Still Lags

None of this matters if the tooling can’t load the model, and K3’s new attention design, Kimi Delta Attention plus Attention Residuals, isn’t something existing inference engines already know how to run.

  • ✓ Unsloth has already published Kimi-K3-GGUF weights across several quant sizes, following the same dynamic-quant methodology used for K2.
  • ✓ REAP-style expert pruning has working precedent on K3's own architecture family, Kimi-Linear and Kimi K2, even without a K3-specific checkpoint yet.
  • × Mainline llama.cpp doesn't recognize K3's architecture yet. Running the GGUF builds currently requires a community fork instead of the standard release.
  • × No confirmed REAP checkpoint or distilled student model built from K3 has shipped as of this writing. What exists so far is a still-enormous quantized original, plus precedent that the rest is coming.

Where You've Met This Already

Release Technique What changed Status
Kimi-K2-Instruct-quantized.w4a16 Post-training quantization BF16 → INT4 weights, ~75% smaller on disk Shipped (K2)
Kimi-Linear-REAP-35B-A3B Expert pruning (REAP) 48B → 35B, 30% of experts removed Shipped (K2 family)
Kimi-K2.6-519B-NVFP4 (REAP keep192) Expert pruning + quantization 50% of experts removed, on top of NVFP4 Shipped (community)
Kimi-K3-GGUF Dynamic quant packaging Native MXFP4 repackaged as GGUF quant ladder Shipped, needs a llama.cpp fork
REAP-pruned Kimi K3 Expert pruning Projected, based on K2 precedent Not yet shipped
Distilled K3 student (14B–30B) Knowledge distillation New, genuinely small dense model Not yet shipped