Vol. IV · 24
Policy Gradients and PPO: Learning Behavior Directly
Reinforcement Learning · Part 3
DQN taught me one way to act intelligently: learn the value of every action, then pick the best one. But there's a second lineage in reinforcement learning with the opposite philosophy — skip the values, and optimize the behavior itself.
Read entry