Skip to main content

Distributional DQN (C51)

Key Insight​

Ordinary DQN predicts a single number per action: the expected return. Distributional RL predicts the whole distribution of possible returns instead — capturing that an action might usually pay off modestly but occasionally win or lose big — which gives the network a far richer training signal and tends to stabilize learning even when you still ultimately act on the mean. C51, the original distributional agent, represents that distribution as a fixed comb of 51 evenly spaced return values (the "atoms") and learns a softmax probability for each; the "51" is simply the atom count that was found to work well — enough resolution to capture the shape of the distribution without so many bins that each is starved of training data. Each Bellman update shifts every atom by the reward and discount factor, redistributes the probabilities back onto the fixed comb (a step called the projection), and trains the network to match that shifted target with a cross-entropy loss.