I Trained a Board-Game AI From Zero Using Self-Play (and Broke It Several Times First)

How Caro5's bot went from a frozen laptop and a useless first model to a generate → train → arena → promote loop running on Modal.

Good Omens Studio5 min readAI Systems

When I started building Caro5 for the Build Small hackathon, I had barely any machine learning knowledge. By the end, I had a self-play training pipeline, an in-browser ONNX model playing at a level that surprised me, and a long list of mistakes I’ll never make again.

This post is about the training loop — the part of the project I’m most proud of, and the part that failed the most.

The starting point: minimax and a frozen laptop

Caro5 is a Caro (gomoku-like) game. Before any neural network existed, the bot was a basic alpha-beta minimax. That was the seed: I put the minimax engine into self-play to generate games, because a model has to learn from something, and I had no dataset.

The first wall was hardware. On my laptop, I could generate a few hundred games before everything froze and nothing else got done. A few hundred games is nowhere near enough to train on, and a frozen laptop means zero iteration speed. Iteration speed, I learned quickly, is the whole game.

Moving generation, training, and evaluation to Modal changed everything. Suddenly game generation was a remote job instead of a machine-killer, and I could build tooling around it: a dashboard where I can filter datasets by generation and schema, download and merge remote datasets, audit data quality, and promote models. If you’re doing any self-play work on consumer hardware, offloading the loop is not an optimization — it’s the difference between having a project and not having one.

The loop

The pipeline that eventually worked:

  1. Generate games with the current best model (merged with games from the previous best, so the dataset doesn’t collapse to one model’s blind spots)
  2. Prepare tensors from the game records
  3. Train a new generation of model
  4. Arena the new candidates against the current champion
  5. Promote the winner if it beats the champion
  6. Repeat

The arena step is the one I’d emphasize to anyone building something similar. Without a head-to-head evaluation gate, you’re just hoping each training run helps. With it, promotion is a measured decision: the new model either beats the incumbent or it doesn’t ship. It’s continuous deployment for models, and it kept me from regressing without noticing.

The first model was useless — and that was a data problem and a me problem

My first trained model simply didn’t work well. Part of it was that I had no idea how to train a model back then. But redoing it taught me the less obvious lesson: most of my problems were data problems wearing a training-problem costume.

Things that mattered far more than I expected:

  • Data quality over data volume. Thousands of games from a weak minimax player teach the model to be a weak minimax player. Merging in games from progressively stronger champions is what moved the ceiling.
  • Schema design is a long-term commitment. My early games were recorded with a minimal schema. Later, when I wanted to train commentary models, I needed rich per-move information — evaluations, threat lists, the model’s “thoughts.” That meant growing the schema to v4, and it meant thousands of early games had zero annotations and were effectively worthless for the new task. Design your data schema for the model you’ll want in three months, not the one you’re training today.

Deployment: the model plays in your browser

The final bot doesn’t call a server for moves. It runs as an ONNX model in the browser, and the move choice is a weighted average of the neural network’s output and a minimax search. Difficulty levels scale the search: Normal runs depth 2 with 16 simulations, and the strongest AI+ mode runs depth 4 with 400 simulations.

One production detail I’m glad I added: a panic mode. If ONNX gets stuck or the bot gets confused, it falls back to a plain minimax move. The user never sees a hang. Hybrid systems with deterministic fallbacks beat pure-ML systems on reliability, and reliability is what a player actually experiences.

What I’d tell past me

  • Get off your laptop immediately; the loop must run remotely or it doesn’t run.
  • Build the arena gate before you build anything fancy. Measurement first.
  • Record more per-example metadata than you think you need. Storage is cheap; re-generating annotated data is not.
  • A weak model that ships behind a fallback beats a strong model that hangs.

The model is still nowhere near the strongest engines out there — I study their strategies and mimic what I can. But it plays at a level I couldn’t have imagined when the laptop was freezing on game generation, and every improvement now flows through a pipeline I trust.

Caro5 is playable here. The training datasets and models are on my Hugging Face profile, and the earlier devlog posts cover days 1–6 of the build.