Policy Evaluation by Matrix Inverse
Key Insight
When the policy is held fixed, the Bellman equation stops being scary: it becomes an ordinary set of linear equations, one per state, that you can solve in a single shot with the matrix inverse V = (I − γPπ)⁻¹ rπ. This policy evaluation step — computing the value function of a given policy — is the easy half of RL; the hard half is improving the policy afterwards. Solving it once by matrix inverse and again by repeated Bellman backups shows that the slow iterative method everyone uses in practice is simply converging to this exact closed-form answer.