PETS / Random Shooting MPC
Key Insight
PETS (Probabilistic Ensembles with Trajectory Sampling) is the simplest serious model-based RL recipe: learn a one-step dynamics model from real transitions, then at every step choose an action by random shooting — sample many random action sequences, roll each one forward through the model (model predictive control), and execute the first action of whichever sequence scored the highest predicted return. Because a single dynamics network is overconfident where it has seen little data, PETS trains an ensemble of probabilistic networks and averages their predictions, so the planner trusts the model only where the ensemble agrees. On a task as small as Pendulum it can match SAC's final performance using a fraction of the environment samples — the headline promise of sample efficiency — because every real transition teaches the model, not just the policy.