DDPG on Pendulum
Key Insight
DDPG (Deep Deterministic Policy Gradient) carries DQN's off-policy, replay-buffer recipe over to continuous control, where max over all actions is no longer a lookup you can compute. It replaces that impossible max with a learned deterministic actor that outputs the single best action directly, while a critic scores it — the deterministic policy gradient flows the critic's gradient back into the actor through the chain rule. Pendulum — swing a single pole upright with a continuous torque and hold it there — is the smallest task where you can watch DDPG learn, and also watch it wobble: with one actor, one critic, and no fix for overestimation, its learning curve is famously fragile, which is exactly the instability TD3 and SAC were built to cure.