Skip to main content

TD-MPC2 Study

Key Insight​

TD-MPC2 plans with a learned dynamics model over a short horizon and then bootstraps from a learned value function beyond it — fusing the strengths of planning (precise short-term decisions) and value learning (cheap long-term foresight), with the Cross-Entropy Method doing the short-horizon action search. Crucially it plans in a learned latent space rather than over raw observations, so the model only has to be accurate about the features the value and policy actually use. With one fixed set of hyperparameters it masters the whole DeepMind Control Suite of continuous-control tasks, making it the current frontier of model-based RL and a clean demonstration of why sharing representations between the model, value, and policy is what finally made model-based control competitive.