Skip to main content

Reparameterization Audit

Key Insight​

SAC's actor outputs a Gaussian that is then squeezed through a tanh squashing function so every action stays inside the environment's bounds, and it trains that actor with the reparameterization trick — sampling plain noise and pushing it through mean + std · ε so gradients flow through the otherwise-random action. The catch is that squashing a Gaussian through tanh changes its probability density, so the log-probability used in the loss needs a correction term; get that correction wrong and SAC keeps running while silently learning the wrong entropy, producing a policy that looks trained but quietly underperforms. Auditing this one formula by hand is the cheapest way to catch a bug that no error message will ever report.