Atari Pong
Key Insight
Moving from CartPole's four numbers to Pong's raw screen pixels forces two classic preprocessing tricks. Frame stacking feeds the network the last four frames at once instead of a single image, because one still frame cannot reveal which way the ball is moving — velocity is only visible across time, and without it the Markov property the algorithm assumes is broken. Reward clipping squashes every score change to −1, 0, or +1 so that games with wildly different point scales can all be trained with one shared set of hyperparameters. With a convolutional Q-network reading the stacked frames, the same algorithm that balanced a pole now learns to beat the built-in Pong opponent from pixels alone — the Atari result that put deep reinforcement learning on the map.