Skip to main content

Mini MBPO

Key Insight​

MBPO (Model-Based Policy Optimization) is the Dyna idea done carefully: train a dynamics model, use it to generate short synthetic rollouts, and mix those fake transitions into the replay buffer that an off-policy learner like SAC trains on. The key design choice is keeping the model rollouts very short — often a single step — because model error compounds with every imagined step, so a short rollout branched from a real state stays trustworthy while still multiplying the data the policy learns from. The payoff is a large sample-efficiency win: the agent reaches a good policy in far fewer real environment steps than a purely model-free baseline, which is exactly what model-based RL promises.