Skip to main content

CEM-MPC (Cross-Entropy Method Model Predictive Control)

Key Insight​

Cross-Entropy Method Model Predictive Control (CEM-MPC) swaps random shooting's blind guessing for the Cross-Entropy Method, an iterative search that fits a distribution to the best action sequences and resamples from it. Each planning step it samples a batch of action sequences, keeps the top-scoring fraction (the "elites"), refits a Gaussian to those elites, and repeats a few times — so the search concentrates around promising actions instead of spraying uniformly, which finds far better plans for the same dynamics model at the cost of more compute per decision. This Cross-Entropy Method shares both its name and its core idea — matching a sampling distribution to a target by minimizing a cross-entropy-like gap — with the cross-entropy loss used to train classifiers, but here the "target" is the set of high-return action sequences rather than the correct labels. CEM is the action search inside Dreamer-era planners and TD-MPC2, which is why it is worth implementing once by hand.