Skip to main content

Count-Based on a Small Env

Key Insight​

Count-based exploration turns "go somewhere new" into a concrete bonus: keep a tally N(s) of how many times each state has been visited and add an intrinsic reward of 1/√N(s) on top of the environment's own reward. Rarely seen states get a large bonus and heavily visited ones get almost none, so a Q-learning agent is actively pulled toward the unexplored frontier instead of relying on the random luck of ε-greedy. Why the 1/√N shape: it mirrors how statistical confidence tightens with more samples, so the bonus fades at exactly the rate your uncertainty about a state does. This is the cleanest illustration of intrinsic motivation — exploration driven by the agent's own curiosity signal — and it works beautifully in small discrete worlds, though it breaks down in large or pixel-based ones where no two observations are ever exactly identical, so every raw count stays stuck at one.