Moonshot AI just shipped the biggest open-weight model anyone’s ever released, and it’s not just big for the sake of being big — there’s real engineering under the hood. Let’s get into what K3 actually does differently, and why it matters even if you never touch the weights yourself.
The headline number isn’t the interesting part
Yes, 2.8 trillion parameters. Yes, first open model in the 3-trillion-parameter class. That’s the number every headline led with, and it’s real — but parameter count on its own is a bit like quoting a building’s total square footage without mentioning how many of those square feet you can actually walk into.
The number that matters more: for any given request, K3 only activates 16 out of 896 experts. Think of it less as one enormous brain and more as a massive team of specialists where a receptionist routes your question to exactly the few people who need to answer it — everyone else stays off the clock. That’s what makes a 2.8T model something you can actually afford to run, instead of a research curiosity that only exists in a press release. The lab behind it, Moonshot AI, has been quietly methodical about this.Moonshot AIthe Beijing-based lab behind the Kimi assistant
parameters in the whole model, the headline number
routed experts the parameters are organized into
experts activated for any given token
If only ~2% of the model works on each token, why train the other 98% at all?
Because different tokens need different specialists. Code tokens recruit different experts than poetry tokens than math tokens. The full pool is what guarantees the right 16 exist for whatever arrives next. You're not paying to run the whole team; you're paying to have the whole team on the roster.
Two architecture changes doing the heavy lifting
K3 is built on two new pieces: Kimi Delta Attention (KDA) and Attention Residuals. Both are aimed at the same underlying problem — as models get deeper and context windows get longer, information has to travel further to get where it’s needed, and it tends to degrade on the way. KDA is a hybrid linear attention mechanism that changes how the model tracks relationships across a sequence; Attention Residuals give information more direct paths through the model’s depth, instead of forcing everything through every layer in sequence.
If linear attention is so much cheaper, why keep any full attention layers at all?
Because a running summary blurs the kind of sharp, needle-in-a-haystack recall that full attention is excellent at. The hybrid keeps linear layers for the long haul and full attention for the moments where exact retrieval matters. Cheap most of the time, precise when it counts.
The practical payoff is a 1-million-token context window that doesn’t fall over under its own weight — plus native visual understanding baked in, not bolted on.
from architecture to training
Quantization wasn’t an afterthought
Most models get quantized after the fact — train it at full precision, then shrink it down for cheaper serving, and eat some quality loss as the cost of doing business. K3 did quantization-aware training starting from the supervised fine-tuning stage, so the model was learning under those constraints from early on rather than getting compressed after the fact.supervised fine-tuningthe stage where a pretrained model is taught to follow instructions from labelled examples
That’s a real distinction, not a technicality.
Post-training quantization
Train at full precision, shrink afterwards, eat the quality loss. Hemming a suit that was cut for someone else.
Quantization-aware training
The model learns under low-precision constraints from the SFT stage onward. Tailoring the suit to a specific body from the first fitting.
It also shows up in the serving math. A 2.8T model at full 16-bit precision is a memory footprint most data centres would feel in their bones; at 4 bits it enters the realm of the merely enormous.4-bit weightseach parameter stored in 4 bits instead of 16, a quarter of the FP16 footprint
The Nook of Wonder Theorems & Beautiful Patterns — what 4 bits does to 2.8 trillion
2.8×10¹² params × 2 bytes (FP16) = 5.6 TB
2.8×10¹² params × 0.5 bytes (INT4) = 1.4 TB
One byte is 8 bits, so FP16 costs 2 bytes per parameter and 4-bit costs half a byte. The same model, a quarter of the memory, before any quality discussion starts.
16 experts / 896 experts = 1/56 ≈ 1.8%
And the routing arithmetic from the first section: roughly 1.8% of the expert pool wakes up per token. Memory you don't load and compute you don't run are the two cheapest things in this industry.
It’s not just competitive for an open model — it’s competitive, period
This is the part worth sitting with: K3 lands around #2–3 on independent benchmark indexes like Vals AI and Artificial Analysis, trailing only a small handful of the very best closed models while costing meaningfully less to run.Vals AIan independent lab that evaluates frontier models on real professional workArtificial Analysisan independent index ranking models on quality, speed, and price It also tops the Frontend Code Arena leaderboard. That’s not “impressive for open-source.” That’s just impressive.
For deployment, Moonshot recommends running it on supernode setups with 64 or more accelerators — a reminder that “open weights” doesn’t mean “runs on your laptop.” Openness here means anyone with the infrastructure can host, fine-tune, and inspect the model, not that the barrier to entry has vanished.
- Anyone with the infrastructure can host the model themselves.
- Anyone can fine-tune the weights for their own use.
- Anyone can inspect exactly what was shipped.
- Not that it runs on your laptop — "open weights" never meant that, and K3 least of all.
the pattern
Where you’ve met this shape before
| Pattern | Earlier example | K3’s version |
|---|---|---|
| Sparse “team of experts” routing | GPT-4 era MoE rumors, Mixtral | Stable LatentMoE, 16-of-896 activation |
| Training-time efficiency baked in | DeepSeek’s cost-efficient training runs | Quantization-aware training from the SFT stage |
| A Chinese lab forcing the ecosystem’s hand | DeepSeek R1’s reasoning-model sprint | K3 pushing Alibaba to open-source its largest Qwen model |