Skip to main content

Eligibility Traces

Key Insight​

Eligibility traces give every recently visited state a fading "credit tag," so that when a TD error finally arrives it updates not just the current state but all the states that led up to it, each in proportion to how recently it was seen. The trace-decay parameter λ slides smoothly between two extreme ways of learning. Setting λ = 0 gives plain one-step TD(0) (updating only the immediate last step), and λ = 1 recovers Monte Carlo (updating every step in the whole episode equally), with intermediate values usually learning fastest. Replacing traces cap a revisited state's trace at 1 rather than adding to it, which stops a state visited in a loop from earning unrealistically large credit. Picture a fading scent trail behind you: when reward finally appears, the freshest parts of the trail feel the strongest pull, but replacing traces ensures a spot you walked in a circle over doesn't smell overpoweringly strong, just "fresh." Sweeping λ ∈ {0, 0.5, 0.9, 1.0} lets you watch the bias–variance dial turn in real time.