RND on Atari
Key Insight
Random Network Distillation (RND) scales the "reward novelty" idea up to high-dimensional Atari screens, where simple visit counts are meaningless because no two frames are ever identical. It keeps two networks: a fixed, randomly-initialized target network and a predictor network trained to copy the target's output on every state the agent visits (distillation of one network into another). On familiar states the predictor has had plenty of practice and its error is tiny; on a genuinely new screen it has never trained, so its error is large — and that prediction error is handed straight to the agent as an intrinsic reward. Why it matters: with no hand-designed counts, RND was the method that finally cracked Montezuma's Revenge, the notorious sparse-reward game where the agent must cross several rooms before earning a single point, putting curiosity-style intrinsic motivation firmly on the map.