First-Visit Monte Carlo
Key Insight
Monte Carlo value estimation throws away the need to know the environment's dynamics: instead of computing expected returns from a model, you simply play many full episodes and average the actual returns that followed each state. The first-visit variant counts, within each episode, only the first time a state is reached — which keeps the averaged samples independent and gives an unbiased estimate of the policy's value function V^π. Think of reviewing a restaurant: if you order the same burger three times in one meal, "first-visit" means your review only counts the very first bite to judge the burger, ignoring the rest so you don't over-count a single experience. Blackjack is the ideal first environment because its rules make the true value of some states easy to work out by hand, so you can check your sampled estimate against an analytic answer — and feel directly how Monte Carlo estimates are unbiased yet jump around a lot until you have averaged thousands of episodes.