Double + Dueling
Key Insight
Standard DQN systematically over-estimates action-values because the same network both picks the best next action and judges its value, so any random noise that makes one action look too good gets trusted — the max always reaches for the luckiest overestimate. Double DQN fixes this with essentially one line: use the online network to choose the next action but the target network to score it, so a fluke in one is unlikely to be echoed by the other. Dueling DQN changes the network's shape instead, splitting it into two streams that separately estimate the state's overall value V(s) and each action's advantage A(s, a) before recombining them as Q = V + A; this lets the agent learn that a state is good or bad without having to try every action there first. Ablating each one on Pong or Breakout shows they are small, independent wins that stack on top of each other.