The Model Already Knew the Answer — It Just Couldn't Stop Talking Long Enough to Say It
A look at a provocative new finding: two random vectors, no training, and a small open-weight model's accuracy nearly doubles.
Read entryFig. I — Filed by category
25 meditations filed under Machine Learning. Back to all categories orthe full archive.
A look at a provocative new finding: two random vectors, no training, and a small open-weight model's accuracy nearly doubles.
Read entryWhy the hardware running today's AI was built to draw video game triangles, and what happens when you point a few thousand of them at the same problem.
Read entryKimi K3 shipped as the largest open-weight model anyone has released. This piece looks at what the compression crowd has actually managed to do about that, two days in.
Read entryBuilding Blocks · Part 13
Article 1 opened with a single weight, multiplying a single number. Nothing since has introduced anything more magical than that.
Read entryBuilding Blocks · Part 9
The exact same matrix multiplications run whether a model is learning or answering. So what could possibly be different enough to need its own dedicated hardware, memory budget, and math?
Read entryBuilding Blocks · Part 6
Before a single vector or tensor can exist, raw text has to be chopped into pieces a model can count on one hand.
Read entryBuilding Blocks · Part 5
Vectors were one row of numbers. Tensors are what you get when a model needs to stack rows into grids, and grids into stacks of grids.
Read entryBuilding Blocks · Part 4
Every entry in an embedding is a separate axis of meaning. Here's what that actually buys a model, and what it costs.
Read entryBuilding Blocks · Part 3
Vectors gave every idea an address. Embeddings are what happens when a model chooses those addresses on purpose, so that meaning has geography.
Read entry