Skip to main content

Policy Evaluation by Matrix Inverse

Key Insight​

When the policy is held fixed, the Bellman equation stops being scary: it becomes an ordinary set of linear equations, one per state, that you can solve in a single shot with the matrix inverse V = (I − γPπ)⁻¹ rπ. This policy evaluation step — computing the value function of a given policy — is the easy half of RL; the hard half is improving the policy afterwards. Solving it once by matrix inverse and again by repeated Bellman backups shows that the slow iterative method everyone uses in practice is simply converging to this exact closed-form answer.