The Bit Explainers · Deep Dive · Open Models(part unknown)

What Kimi K3 Actually Changed

Moonshot AI's 2.8-trillion-parameter model is the largest open-weight release ever shipped. The size is the headline. The engineering underneath is the story.

Good Omens Studio9 min readEngineering

Moonshot AI just shipped the biggest open-weight model anyone’s ever released, and it’s not just big for the sake of being big — there’s real engineering under the hood. Let’s get into what K3 actually does differently, and why it matters even if you never touch the weights yourself.

The headline number isn’t the interesting part

Yes, 2.8 trillion parameters. Yes, first open model in the 3-trillion-parameter class. That’s the number every headline led with, and it’s real — but parameter count on its own is a bit like quoting a building’s total square footage without mentioning how many of those square feet you can actually walk into.

The number that matters more: for any given request, K3 only activates 16 out of 896 experts. Think of it less as one enormous brain and more as a massive team of specialists where a receptionist routes your question to exactly the few people who need to answer it — everyone else stays off the clock. That’s what makes a 2.8T model something you can actually afford to run, instead of a research curiosity that only exists in a press release. The lab behind it, Moonshot AI, has been quietly methodical about this.Moonshot AIthe Beijing-based lab behind the Kimi assistant

TOTAL2.8T

parameters in the whole model, the headline number

SPECIALISTS896

routed experts the parameters are organized into

PER REQUEST16

experts activated for any given token

In one sentenceA mixture-of-experts model holds hundreds of specialist sub-networks and activates only a handful per token, so total capacity and per-request cost come apart.
IntuitionA hospital doesn't route every patient through every specialist. A triage nurse listens to your complaint and sends you to the two people who actually help. The hospital's expertise is enormous; the cost of your visit is small. MoE routing is triage, applied to tokens.
TechnicalA sparse MoE layer replaces one dense feed-forward block with many expert blocks plus a gating network. The gate scores every expert for the current token, selects the top-k (here k = 16 of 896, under Moonshot's Stable LatentMoE variant), and sums only those outputs. Attention layers stay shared; the routed experts carry most of the parameter count, which is why active parameters per token are a small fraction of the total.

If only ~2% of the model works on each token, why train the other 98% at all?

Because different tokens need different specialists. Code tokens recruit different experts than poetry tokens than math tokens. The full pool is what guarantees the right 16 exist for whatever arrives next. You're not paying to run the whole team; you're paying to have the whole team on the roster.

Two architecture changes doing the heavy lifting

K3 is built on two new pieces: Kimi Delta Attention (KDA) and Attention Residuals. Both are aimed at the same underlying problem — as models get deeper and context windows get longer, information has to travel further to get where it’s needed, and it tends to degrade on the way. KDA is a hybrid linear attention mechanism that changes how the model tracks relationships across a sequence; Attention Residuals give information more direct paths through the model’s depth, instead of forcing everything through every layer in sequence.

In one sentenceKimi Delta Attention compresses what it has read into a running state, so the cost of reading grows with the sequence instead of with its square.
IntuitionStandard attention re-reads the whole book every time it turns a page, comparing the new sentence against every previous one. Linear attention keeps a running summary in the margin instead: update the summary as you read, consult the summary when you need to. Much cheaper per page, and the summary is the trick.
TechnicalLinear attention replaces the softmax over all key-value pairs with a recurrent state: keys and values are folded into a fixed-size matrix updated per token, giving linear time per step instead of quadratic. KDA adds a delta-rule update, overwriting stale information in the state rather than letting it accumulate as noise. K3 interleaves these linear layers with occasional full-attention layers.

If linear attention is so much cheaper, why keep any full attention layers at all?

Because a running summary blurs the kind of sharp, needle-in-a-haystack recall that full attention is excellent at. The hybrid keeps linear layers for the long haul and full attention for the moments where exact retrieval matters. Cheap most of the time, precise when it counts.

The practical payoff is a 1-million-token context window that doesn’t fall over under its own weight — plus native visual understanding baked in, not bolted on.

standard residual stack all traffic through every layer + attention residuals direct paths skip the queue
Attention Residuals, schematically: early layers keep a direct line to much deeper ones, so signal doesn't have to survive every intermediate transformation.

from architecture to training

Quantization wasn’t an afterthought

Most models get quantized after the fact — train it at full precision, then shrink it down for cheaper serving, and eat some quality loss as the cost of doing business. K3 did quantization-aware training starting from the supervised fine-tuning stage, so the model was learning under those constraints from early on rather than getting compressed after the fact.supervised fine-tuningthe stage where a pretrained model is taught to follow instructions from labelled examples

That’s a real distinction, not a technicality.

Post-training quantization

Train at full precision, shrink afterwards, eat the quality loss. Hemming a suit that was cut for someone else.

Quantization-aware training

The model learns under low-precision constraints from the SFT stage onward. Tailoring the suit to a specific body from the first fitting.

It also shows up in the serving math. A 2.8T model at full 16-bit precision is a memory footprint most data centres would feel in their bones; at 4 bits it enters the realm of the merely enormous.4-bit weightseach parameter stored in 4 bits instead of 16, a quarter of the FP16 footprint

The Nook of Wonder Theorems & Beautiful Patterns — what 4 bits does to 2.8 trillion

2.8×10¹² params × 2 bytes (FP16) = 5.6 TB
2.8×10¹² params × 0.5 bytes (INT4) = 1.4 TB

One byte is 8 bits, so FP16 costs 2 bytes per parameter and 4-bit costs half a byte. The same model, a quarter of the memory, before any quality discussion starts.

16 experts / 896 experts = 1/56 ≈ 1.8%

And the routing arithmetic from the first section: roughly 1.8% of the expert pool wakes up per token. Memory you don't load and compute you don't run are the two cheapest things in this industry.

It’s not just competitive for an open model — it’s competitive, period

This is the part worth sitting with: K3 lands around #2–3 on independent benchmark indexes like Vals AI and Artificial Analysis, trailing only a small handful of the very best closed models while costing meaningfully less to run.Vals AIan independent lab that evaluates frontier models on real professional workArtificial Analysisan independent index ranking models on quality, speed, and price It also tops the Frontend Code Arena leaderboard. That’s not “impressive for open-source.” That’s just impressive.

For deployment, Moonshot recommends running it on supernode setups with 64 or more accelerators — a reminder that “open weights” doesn’t mean “runs on your laptop.” Openness here means anyone with the infrastructure can host, fine-tune, and inspect the model, not that the barrier to entry has vanished.

  • Anyone with the infrastructure can host the model themselves.
  • Anyone can fine-tune the weights for their own use.
  • Anyone can inspect exactly what was shipped.
  • Not that it runs on your laptop — "open weights" never meant that, and K3 least of all.

the pattern

Where you’ve met this shape before

Pattern Earlier example K3’s version
Sparse “team of experts” routing GPT-4 era MoE rumors, Mixtral Stable LatentMoE, 16-of-896 activation
Training-time efficiency baked in DeepSeek’s cost-efficient training runs Quantization-aware training from the SFT stage
A Chinese lab forcing the ecosystem’s hand DeepSeek R1’s reasoning-model sprint K3 pushing Alibaba to open-source its largest Qwen model