Q-Learning on FrozenLake
Key Insight
Q-learning is the first algorithm in this guide that learns purely from experience without a model (meaning it doesn't know the rules of the world beforehand, like a child learning to ride a bike by falling and balancing rather than reading a physics textbook): it watches (state, action, reward, next-state) transitions and nudges its action-value estimate Q(s, a) toward "reward now plus the discounted value of the best next action." That "best next action" inside the target — rather than the action actually taken — is what makes Q-learning off-policy: it learns the optimal policy while still wandering randomly through ε-greedy action selection, with ε decayed over time so the agent explores early and exploits later. On slippery FrozenLake, tracking the final greedy policy's success rate shows tabular Q-learning recovering a near-optimal policy from nothing but sampled rewards.