Skip to main content

PPO on Atari

Key Insight​

Moving PPO from low-dimensional control to Atari tests whether the same on-policy recipe survives raw pixel input, and the answer is yes — with the standard image-RL scaffolding bolted on. The policy and value networks now share a convolutional trunk that reads the screen, the last few frames are stacked together (frame stacking) so the agent can perceive motion and direction from otherwise-static images, and rewards are squashed to a fixed range (reward clipping) so games with wildly different score scales train with one set of hyperparameters. Because PPO is on-policy, it is data-hungry and leans heavily on running many environment copies in parallel to feed each update. Comparing your agent's scores on a few games against published numbers is the honest check that every detail is wired correctly.