The Model Already Knew the Answer — It Just Couldn't Stop Talking Long Enough to Say It
A look at a provocative new finding: two random vectors, no training, and a small open-weight model's accuracy nearly doubles.
Read entryFig. I — Filtered by tag
6 meditations tagged inference. Back toall tags or the full archive.
A look at a provocative new finding: two random vectors, no training, and a small open-weight model's accuracy nearly doubles.
Read entryKimi K3 shipped as the largest open-weight model anyone has released. This piece looks at what the compression crowd has actually managed to do about that, two days in.
Read entryBuilding Blocks · Part 9
The exact same matrix multiplications run whether a model is learning or answering. So what could possibly be different enough to need its own dedicated hardware, memory budget, and math?
Read entryTraining a capable model is only half the battle. The model that comes out of pretraining or fine-tuning is almost never the model you actually want to ship.
Read entryQuantization · Part 1
Every series I've written so far has bumped into quantization from the side — QLoRA compressing frozen bases, the optimization guide's memory math, GGUF files on Hugging Face with cryptic suffixes.
Read entryLoRA Deep Dive · Part 4
The first three articles in this series were about making an adapter. This one is about what makes adapters a genuinely different kind of artifact from a fine-tuned model: what you can do with them afterward.
Read entry