Skip to main content

ICM

Key Insight​

The Intrinsic Curiosity Module (ICM) builds its novelty bonus from a forward model — a network that predicts the next state from the current state and action — and pays the agent an intrinsic reward equal to how wrong that prediction was, on the logic that surprising transitions are worth revisiting. Its key refinement over RND is where it measures surprise: ICM first learns a compact feature space using an inverse model (predict which action took you from one state to the next), which keeps only the parts of an observation the agent can actually control and discards uncontrollable visual detail. Why that matters: predicting in this controllable feature space partly protects ICM from the noisy-TV problem, where a raw pixel-level predictor would chase random flicker forever. Running ICM and RND on the same task lays bare two different bets on one idea — model the dynamics versus distill a random function — both aimed at turning prediction error into directed exploration.