Reinforcement Learning · Part 1

Reinforcement Learning from Zero: Agents, Rewards, and the Loop That Learns from Consequences

Supervised learning felt intuitive to me from day one: here's the input, here's the right answer, minimize the difference.

Good Omens Studio7 min readReinforcement Learning

Supervised learning felt intuitive to me from day one: here’s the input, here’s the right answer, minimize the difference. Reinforcement learning broke my brain a little, because it removes the answer key entirely. Nobody tells the agent what the right action was. It acts, the world responds, and occasionally a number arrives — a reward — that says “that went well” or “that went badly,” often long after the action that actually mattered. Learning under those conditions is a fundamentally harder, stranger, and more interesting problem, and it’s the paradigm behind everything from game-playing superhuman AIs to the alignment of chatbots. This post lays the foundation.

The loop

All of RL is one loop. An agent observes the state of an environment, chooses an action, and the environment returns two things: a new state and a reward — a single scalar. Repeat until the episode ends. The agent’s goal is not to get the biggest reward right now but to maximize the return: the total accumulated reward over time.

That’s it. Chess: state = board, action = move, reward = +1 for winning, 0 the rest of the game. Robot: state = joint angles and camera pixels, action = motor torques, reward = forward progress. A dialogue model: state = the conversation so far, action = the next token, reward = a human’s approval at the end. The generality of the framing is the point — almost any goal-directed problem can be poured into this mold.

The formal skeleton is the Markov Decision Process (MDP): states, actions, transition probabilities, rewards, and the Markov assumption — the current state summarizes everything you need; history adds nothing. Real problems violate this constantly (a single camera frame doesn’t tell you velocity), which we patch by stacking recent observations into the state or by giving the agent memory. What the agent learns is a policy, written π(a|s): a mapping from states to (possibly probabilistic) actions. The policy is the behavior. Everything in RL is ultimately in service of improving π.

Two ideas that make it hard — and interesting

Delayed credit. You win a chess game on move 60, and the reward arrives. Which of your 60 moves deserve the credit? Maybe the game was decided by a quiet pawn move on move 14. This is the credit assignment problem, and it’s the central difficulty of RL: rewards are sparse, delayed, and unlabeled with causes. Most of the field’s machinery exists to smear that terminal reward backward through time onto the decisions that earned it.

The standard first tool is discounting: future rewards are multiplied by γ^t (gamma slightly below 1, like 0.99), so reward arriving t steps from now is worth γ^t as much as reward now. Partly this reflects genuine impatience and uncertainty about the future; mathematically, it keeps infinite sums finite and, practically, it sets the agent’s planning horizon — γ=0.99 means the agent effectively cares about the next few hundred steps.

Exploration vs. exploitation. A supervised model never has to decide what data to see — the dataset is given. An RL agent generates its own data by acting, which creates a genuine dilemma: keep doing what has worked (exploit) or try something new that might work better (explore)? Pure exploitation gets you stuck ordering the same decent dish at the same restaurant forever, never discovering the better one two doors down; pure exploration never cashes in on anything learned. Simple recipes — act randomly with probability ε, or sample actions in proportion to their estimated quality, or add an entropy bonus that rewards indecision itself — are shockingly load-bearing in practice. And because the agent’s data depends on its current policy, RL data is non-stationary and self-generated: a bad policy explores bad states and learns about nothing else. This feedback loop between behavior and data is what makes RL so much less stable than supervised learning.

Value functions: learning what states are worth

The conceptual centerpiece of classical RL is the value function. V(s) asks: starting from state s and following my policy, what total return should I expect? Its sibling Q(s, a) asks the more actionable version: what return should I expect if I take action a in state s, then follow my policy? If you knew Q perfectly, acting optimally would be trivial — in every state, take the action with the highest Q. No planning, no search, just lookup.

Values obey a beautiful recursive consistency called the Bellman equation: the value of a state equals the immediate reward plus the discounted value of the next state. Today’s worth is today’s income plus tomorrow’s worth. This turns the impossible-sounding problem of “estimate total future reward” into a local bootstrapping problem: make your estimate for this state consistent with your estimate for the next state, everywhere, and the true values fall out.

Temporal-difference (TD) learning operationalizes it: after each step, compare what you predicted (V(s)) with what you just observed (reward + γV(s′)), and nudge the prediction toward that target. The gap between them — the TD error — is the “surprise” signal that drives learning: things went better than expected, value up; worse, value down. One of my favorite facts in all of science: dopamine neurons in mammalian brains fire in a pattern that closely matches the TD error. The algorithm was derived from math; evolution apparently derived it first.

Q-learning applies the same trick to Q with one bold twist: bootstrap not from the action you actually took next, but from the best action available (the max over Q(s′, ·)). This makes it off-policy — you can behave sloppily (exploring, even acting randomly) while still learning the values of the optimal policy. Given enough exploration, a table of Q-values provably converges to optimal. That guarantee holds for small, table-sized worlds — and breaks down exactly at the point where things get interesting, when states are too numerous to enumerate and we bring in neural networks. That collision (and its resolution, DQN) is the next article.

The map of the territory

Before going deeper, it helps to see the family tree, because every RL method is a stance on a few dichotomies:

  • Value-based vs. policy-based. Learn Q and derive actions from it (DQN lineage) — or skip values and directly adjust the policy’s parameters toward higher reward (policy gradients, PPO — article 3). Actor-critic methods, the modern default, do both: a policy (“actor”) learns to act, a value function (“critic”) learns to judge, each improving the other.
  • On-policy vs. off-policy. Must the training data come from the current policy (fresh but wasteful) or can it come from anywhere — old policies, replay buffers, other agents’ demonstrations (efficient but trickier to keep stable)?
  • Model-free vs. model-based. Learn behavior directly from experience — or learn a model of the environment and plan inside it, imagining futures before living them. Model-based methods (the MuZero lineage) are far more sample-efficient when the model is good and dangerously deluded when it isn’t.

Why this framework conquered more than games

For years RL’s showcase was games — Atari, Go, StarCraft — and it was fair to ask whether it was anything more than a very expensive way to win at them. The answer arrived from an unexpected direction: language models. The problem of making an LLM helpful and harmless is, structurally, an RL problem — the “right answer” isn’t defined per token, only a fuzzy judgment of whether the whole response was good. That judgment is a reward. RLHF (article 4) is TD-era ideas wearing a new coat, and the recent wave of reasoning models — trained with RL against verifiable rewards like “did the math check out” — pushed it further: the models discovered longer chains of thought because thinking longer earned more reward. Nobody labeled the reasoning steps. Consequences taught them.

That, ultimately, is the thing RL contributes that supervised learning cannot: the ability to exceed the teacher. A supervised model is capped by its labels — it can only imitate. An agent trained on consequences can find strategies no human demonstrated, from AlphaGo’s move 37 to proof strategies and code no dataset contained. Learning from an answer key gets you to human level. Learning from consequences is how you pass it.