TD3 on HalfCheetah
Key Insight
TD3 (Twin Delayed DDPG) keeps DDPG's actor-critic skeleton but adds three fixes that turn a fragile algorithm into a dependable one: twin critics whose smaller estimate is used as the target, so the policy cannot exploit one critic's lucky overestimate; delayed policy updates that let the critic settle before the actor chases it; and target policy smoothing, a little noise added to the target action so the critic cannot overfit to a razor-thin peak. HalfCheetah — a two-legged running robot simulated in MuJoCo — is the standard benchmark where these fixes visibly lift TD3's returns above DDPG's noisy, often-diverging ones.