DQN on CartPole
Key Insight
DQN replaces Q-learning's lookup table with a small neural network (an MLP) that maps a state to one action-value per action — this is function approximation, and it is what lets the same algorithm scale from a 16-square grid to environments with far too many states to ever tabulate. CartPole is the gentlest place to watch this work: the state is just four numbers (cart position, cart velocity, pole angle, pole angular velocity), so a tiny network can learn to balance the pole in under 30,000 steps. Stripping out the usual stabilizers — no experience replay, no target network — exposes the raw fragility the deadly triad warns about: training often climbs and then suddenly collapses, because the network is chasing a learning target built from its own constantly-shifting predictions. Seeing that instability firsthand is the whole point, and it motivates every fix in the projects that follow.