Skip to main content

Value Iteration on FrozenLake

Key Insight​

FrozenLake is the smallest Gymnasium environment whose dynamics are fully known, which is exactly what value iteration needs: when the transition probabilities and rewards are written down, you can compute the optimal value function V* by planning — repeatedly applying the optimality Bellman backup — instead of learning it from trial and error. Because the ice is slippery, the agent sometimes slides sideways instead of where it aimed, so the optimal policy must respect those random transitions and often points away from the nearest hole rather than straight at the goal. Visualizing V* as a heatmap over the grid and the greedy policy as arrows turns the abstract fixed-point computation into something you can see at a glance.