Skip to main content

Atari Pong

Key Insight​

Moving from CartPole's four numbers to Pong's raw screen pixels forces two classic preprocessing tricks. Frame stacking feeds the network the last four frames at once instead of a single image, because one still frame cannot reveal which way the ball is moving — velocity is only visible across time, and without it the Markov property the algorithm assumes is broken. Reward clipping squashes every score change to −1, 0, or +1 so that games with wildly different point scales can all be trained with one shared set of hyperparameters. With a convolutional Q-network reading the stacked frames, the same algorithm that balanced a pole now learns to beat the built-in Pong opponent from pixels alone — the Atari result that put deep reinforcement learning on the map.