Skip to main content

REINFORCE on CartPole

Key Insight​

REINFORCE is the most direct way to learn a policy: instead of estimating action values and acting greedily, it nudges the policy's weights so that actions taken in successful episodes become more likely and actions from failed ones become less likely. The mechanism is the policy gradient theorem, made computable by the log-derivative trick, which weights each action by the full Monte Carlo return of its rollout. CartPole is the gentlest place to see this work — but also to feel its flaw: because the weight is a whole noisy trajectory's return rather than a one-step estimate, the gradient is unbiased but high-variance, so training lurches around and learns slowly. Watching that variance firsthand is the whole point, and it motivates the baseline and actor-critic fixes in the projects that follow.

REINFORCE vs. DQN​

REINFORCE and DQN (Deep Q-Network) represent the two primary branches of reinforcement learning: policy-gradient methods and value-based methods.

  • DQN (Value-Based): Learns to estimate the expected future reward (value) of taking each action in a given state (using a critic or value network), and then acts greedily based on those estimates. It is off-policy and relies on an experience replay buffer to learn from past data.
  • REINFORCE (Policy-Gradient): Bypasses estimating action values entirely. It directly outputs a probability distribution over actions (the policy) and updates its weights using Monte Carlo returns from complete rollouts. It is strictly on-policy.

Analogy: Imagine learning to play golf.

  • A DQN golfer (value-based) tries to calculate the exact expected score or landing position for every possible club and swing angle, and then picks the one with the highest estimated value.
  • A REINFORCE golfer (policy-gradient) doesn't calculate any values. They just swing. If the ball lands in or near the hole, they try to repeat that exact swing next time. If the ball lands in a pond, they avoid that swing in the future.

While DQN is often more sample-efficient because it can reuse past data, REINFORCE is mathematically simpler, directly optimizes the policy, and works naturally in continuous action spaces where calculating values for infinitely many actions is impossible.