I Trained a Board-Game AI From Zero Using Self-Play (and Broke It Several Times First)
How Caro5's bot went from a frozen laptop and a useless first model to a generate → train → arena → promote loop running on Modal.
Read entryFig. I — Filtered by tag
5 meditations tagged reinforcement learning. Back toall tags or the full archive.
How Caro5's bot went from a frozen laptop and a useless first model to a generate → train → arena → promote loop running on Modal.
Read entryReinforcement Learning · Part 4
A base language model fresh out of pretraining is a strange creature. It has read a large fraction of the internet and can continue any text with uncanny fluency — but it isn't trying to help you.
Read entryReinforcement Learning · Part 3
DQN taught me one way to act intelligently: learn the value of every action, then pick the best one. But there's a second lineage in reinforcement learning with the opposite philosophy — skip the values, and optimize the behavior itself.
Read entryReinforcement Learning · Part 2
In 2013, a small London startup called DeepMind posted a paper showing a single algorithm learning to play Atari games — from raw pixels, with no game-specific knowledge, using only the score as feedback.
Read entryReinforcement Learning · Part 1
Supervised learning felt intuitive to me from day one: here's the input, here's the right answer, minimize the difference.
Read entry