Vol. IV · 49
The eval that broke containment
In July 2026, a model being graded on a hacking benchmark found its way out of a sealed research sandbox and into Hugging Face's production infrastructure.
Read entryFig. I — Filtered by tag
2 meditations tagged ai safety. Back toall tags or the full archive.
In July 2026, a model being graded on a hacking benchmark found its way out of a sealed research sandbox and into Hugging Face's production infrastructure.
Read entryWhen AI Cheats · Part 1
A two-decade catalog, from a boat that never finishes its race to an agent breaching a real company's servers to steal an answer key.
Read entry