DQN taught me one way to act intelligently: learn the value of every action, then pick the best one. But there’s a second lineage in reinforcement learning with the opposite philosophy — skip the values, and optimize the behavior itself. Parameterize a policy with a neural network, measure how much reward it collects, and push its parameters directly toward more reward. This is the policy-gradient family, it handles problems DQN structurally can’t, and its most famous member — PPO — became the default algorithm of applied RL and the engine that aligned ChatGPT-era language models. This is the article where the math got genuinely beautiful for me, so I’ll try to earn each equation-in-words.
Why go direct?
Value-based methods have blind spots. The max over actions breaks on continuous control — you can’t enumerate all possible torque vectors for a robot arm. They commit to deterministic behavior, but some situations genuinely demand stochastic policies (in poker, any predictable bluffing pattern is exploitable; in partially observed worlds, two identical-looking states may need different action mixes). And they optimize an indirect quantity: a slightly better Q-estimate can flip an argmax and lurch the behavior discontinuously.
A policy network sidesteps all of it. The network takes the state and outputs action probabilities (or, for continuous actions, the mean and spread of a distribution to sample from). Behavior changes smoothly as parameters change, stochasticity is built in — exploration is just the policy’s own randomness, annealing naturally as it grows confident — and continuous actions are no harder than discrete ones.
One question remains, and it’s the whole problem: the reward comes from sampling actions and letting the environment respond. You can’t backpropagate through a dice roll and a game of Pong. So how do you get a gradient?
The score function trick: the one idea at the center
The policy-gradient theorem (in its REINFORCE form) answers with a move so simple it feels like a con: make good actions more probable in proportion to how good their outcomes were.
Run the policy, collect trajectories. For each action taken, compute the return that followed it. Then adjust the network to increase the log-probability of each action, scaled by that return. Actions followed by high reward get pushed up; actions followed by poor reward get pushed up less — or down. Averaged over many samples, this is provably an unbiased estimate of the true gradient of expected reward. No model of the environment, no differentiating through game physics — the environment stays a black box, and the randomness of your own policy is what lets you learn from it.
Notice the philosophical difference from supervised learning: nobody says which action was correct. The signal is just “whatever you did there — more of that” weighted by results. It’s trial and error, formalized into calculus.
And in raw form, it barely works. The reason is variance. Returns are noisy — a mediocre action during a lucky episode gets credit; a brilliant action in a doomed one gets blamed. The gradient signal is real but drowning in noise, so raw REINFORCE needs absurd numbers of samples and still trains erratically. The whole subsequent history of policy gradients is variance-reduction engineering, in three steps:
Only credit what followed. An action can’t cause rewards that arrived before it — so scale each action by the reward-to-go from that point on, not the whole episode’s return. Free variance reduction from pure causality.
Subtract a baseline. What matters isn’t whether the return was high, but whether it was higher than expected from that state. Subtract V(s) — expected return from the state — and you get the advantage: A(s,a) = (what happened after taking a) − (what typically happens from s). Scoring 3 points is brilliant from a losing position, disappointing from a winning one; the advantage encodes exactly that. Subtracting a baseline provably doesn’t bias the gradient — it only cuts the noise.
Learn the baseline. Where does V come from? A second network, trained to predict returns. Now you have actor-critic: an actor (the policy) chooses actions; a critic (the value network) judges situations, providing the baseline that makes the actor’s learning signal clean. The critic sees a state and says “this is usually a +5 situation”; the actor gets graded on beating or missing that par. Add the standard refinement GAE — generalized advantage estimation, a knob for smoothly trading the critic’s bias against the sample noise — and you have the modern skeleton: nearly every headline RL system of the last decade is an actor-critic under the hood.
The step-size problem, and trust regions
One more failure mode stands between this and a usable algorithm, and it’s peculiar to RL. In supervised learning, an oversized gradient step wastes an update — the data is still there, you recover. In on-policy RL, the policy collects its own data. Take one bad step, and the degraded policy now explores garbage states, generating garbage data, from which it learns further garbage. Performance doesn’t dip; it collapses, sometimes unrecoverably. The lesson: policy updates must be kept small in behavior, not merely small in parameters — a tiny weight change can be a huge behavior change.
TRPO (2015) made this rigorous: maximize improvement subject to a hard constraint on how much the action distribution may shift (measured by KL divergence), backed by monotonic-improvement theory. It worked, and was miserable to implement — second-order optimization, conjugate gradients, incompatible with everyday tooling.
PPO (2017) is Schulman and colleagues asking: can we get TRPO’s caution with SGD’s simplicity? Their answer is one clipped objective. For each action, compute the probability ratio — how much more (or less) likely the new policy is to take that action than the old policy that collected the data. Multiply by the advantage, as usual. But clip the ratio to a small interval around 1 (canonically ±20%): beyond that band, the objective flat-lines, so the gradient for pushing further vanishes. The incentive to exploit an advantage simply switches off once the policy has moved 20% away from where the data was collected. No constraint solver, no second-order math — a clip function. First-order optimizers, a few epochs of minibatch reuse per data batch (a modest but real sample-efficiency win over vanilla policy gradients), done.
PPO is not the most sample-efficient algorithm, nor the most theoretically airtight — the clip is a heuristic that usually keeps updates in a trust region. What it is, is robust: it works across continuous control, games, and language with minimal tuning, implemented in a page of code. That robustness made it the field’s default — OpenAI Five (Dota 2), dexterous robot hands, and countless production systems run on PPO. In applied RL, the algorithm that reliably works beats the algorithm that occasionally works better.
The twist ending: PPO’s biggest domain isn’t robotics
Here’s what I find delightful. PPO was built for simulated locomotion and Atari — and its most consequential deployment turned out to be language. In RLHF, a language model is a policy: state = the conversation so far, action = the next token, an episode = a full response, reward = a score from a model trained on human preferences. PPO fine-tunes the LLM to maximize that reward — with the clip, plus an extra KL leash to the original model, preventing the policy from collapsing into reward-hacking gibberish. The same “don’t move too far from where your data came from” caution that kept simulated cheetahs from faceplanting is what kept aligned chatbots coherent.
That story — human preferences, reward models, the KL leash, and the newer methods like DPO that shortcut PPO entirely — is the next article. But the through-line of this one is worth restating: policy gradients are the mathematics of “do more of what worked.” Everything else — advantages, critics, clipping — is the hard-won engineering of making that childlike principle stable enough to train something real.