A base language model fresh out of pretraining is a strange creature. It has read a large fraction of the internet and can continue any text with uncanny fluency — but it isn’t trying to help you. Ask it a question and it might answer, or generate three more questions like yours (that’s a plausible continuation too), or wander into a forum flame war it half-remembers. The gap between “predicts text” and “behaves like an assistant” is where reinforcement learning enters the story of LLMs — and closing that gap, via RLHF, is arguably what turned language models from a research curiosity into the product wave we’re living through. This final RL post connects everything from the previous three to the models we actually use.
The core difficulty: the objective can’t be written down
What we want from an assistant — helpful, honest, harmless, appropriately concise, right tone — is easy to recognize and impossible to specify. You cannot write a loss function for “good response.” Supervised learning offers a partial fix: hire people to write ideal responses and fine-tune on them (SFT — supervised fine-tuning). This works remarkably well as step one; it teaches the format of assistance — answer the question, be polite, stop when done.
But SFT alone hits three walls. Demonstrations are expensive and cap the model at imitating its teachers. Imitation can’t express don’t — there’s no way to show “never do this” as a demonstration. And crucially, humans are far better at judging responses than producing them — anyone can tell which of two answers is better; writing the perfect one is hard. That asymmetry is the resource RLHF is built to mine: if people can rank outputs, rankings can become a reward, and RL can maximize it.
The RLHF pipeline
The recipe that InstructGPT established (2022) — and that, scaled up, became ChatGPT — has three stages:
1. SFT. Fine-tune the pretrained model on human-written demonstrations. This produces a competent-but-rough assistant, and its real role in the pipeline is to be a sane starting policy — RL from a raw base model would explore hopelessly.
2. Train a reward model. Sample several responses per prompt, have human labelers rank them, and train a separate network — usually a clone of the LLM with a scalar output head — to predict those preferences: given (prompt, response), emit a score such that preferred responses score higher (via the Bradley–Terry framing: the probability humans prefer A over B is a function of score(A) − score(B)). This is the pivotal move in the whole design: compress diffuse human judgment into a differentiable, queryable function. Human taste, at one remove, becomes something you can optimize against — millions of times, for free, without asking a human again.
3. RL against the reward model. Now it’s the previous article’s machinery, transplanted: the LLM is a policy, the prompt is the initial state, each token is an action, a full response is an episode, and the reward model scores the finale. PPO adjusts the model to produce higher-scoring responses — with one addition that carries most of the safety load: a KL penalty to the SFT model. Every response is penalized in proportion to how far its token probabilities drift from the reference model. The final objective is, in words: maximize reward, minus λ times how weird you’ve become.
Why the leash? Because of the failure mode that defines this whole field: reward hacking. The reward model is not human judgment — it’s a lossy proxy, accurate on the distribution it was trained on and exploitable off it. Optimize hard enough and PPO will find the cracks: responses that are longer (labelers drift toward longer answers), more flattering, more confident-sounding, stuffed with hedges or bullet lists or whatever superficial features correlate with high scores — including fluent nonsense that scores wonderfully. This is Goodhart’s law running at machine speed: when a measure becomes a target, it ceases to be a good measure. The KL term keeps the policy in the region where the proxy still tracks the truth. Loosen it and the model degrades in ways the reward model can’t see; overtighten and nothing improves. Tuning that tension is the craft of RLHF.
It’s worth pausing on what the result means. An RLHF’d model isn’t “trained on more text.” It’s been shaped by an optimization pressure toward human approval — which explains both its virtues (responsiveness, care, refusal of harmful requests) and its characteristic vice: sycophancy, the learned instinct that agreement scores well. Optimizing approval is not identical to optimizing truth, and every assistant you’ve ever used lives somewhere on that tension.
DPO: the shortcut that shook the pipeline
Full RLHF is an engineering ordeal — four models in memory (policy, reference, reward model, critic), PPO’s instability, endless tuning. In 2023, the Direct Preference Optimization paper asked a question with an almost embarrassing answer: is the reward-model-plus-PPO detour necessary at all?
Their derivation: take the exact objective RLHF optimizes (maximize reward under a KL leash). Its optimal solution can be written in closed form — and inverted, so that the reward is expressed in terms of the optimal policy itself. Substitute that back into the preference-model math, and the reward model cancels out of the equation entirely. What remains is a plain classification-style loss, directly on preference pairs: raise the likelihood of preferred responses, lower the likelihood of rejected ones, with each example weighted by an implicit KL leash to the reference model. As the paper’s title put it — your language model is secretly a reward model.
No sampling during training, no PPO, no separate reward network: preferences in, aligned model out, with the stability of supervised fine-tuning. DPO and its descendants (IPO, KTO — which needs only thumbs-up/down rather than pairs, SimPO, ORPO) swept open-source alignment almost overnight.
The honest comparison, though, is not “DPO won.” DPO optimizes against the static dataset of preferences; PPO-style RLHF optimizes against a reward model that generalizes beyond it, and — critically — trains on-policy, on the model’s own fresh outputs rather than whoever generated the preference data. At the frontier, labs largely still run reward-model-based, on-policy RL (with many hybrid variants); DPO’s kingdom is everywhere that four-model PPO infrastructure is too expensive. The practical hierarchy I’ve settled on: SFT teaches format, DPO-family methods cheaply instill preferences, full RL extracts the most from a reward signal — at the most cost and the most Goodhart risk.
The plot twist: rewards you can verify
Everything above optimizes a learned, fuzzy proxy for quality. The most important recent development flips that premise. For some domains, reward needs no human and no proxy: the answer is checkable. Did the math produce the right number? Does the code pass the tests? Did the proof verify? This is RLVR — RL from verifiable rewards — and it’s the training paradigm behind the reasoning-model era (o1, R1, and kin).
The setup is almost austere: hard problems, a checker, and RL (typically simplified PPO variants like GRPO, which drops the critic and baselines each response against other samples for the same prompt). What emerged stunned everyone: models spontaneously learned to think longer — generating extended chains of reasoning, second-guessing themselves, backtracking, verifying — because deliberation earned reward. Nobody demonstrated those behaviors; they were discovered, the way AlphaGo discovered move 37, because a checkable objective plus exploration finds strategies no dataset contained. And a verifiable reward can’t be Goodharted the way a preference model can: the test either passes or it doesn’t.
Which brings the four-article arc to a clean close. RL’s founding promise was learning that exceeds the teacher — from consequences, not labels. For years the showcases were games. It turns out the same loop — policy, reward, credit assignment, and the eternal engineering of stability — now shapes how machines write, reason, and refuse. The reward signals changed: game scores, then human preferences, then verifiable truth. The loop never did.