Skip to main content

A2C with Parallel Environments

Key Insight​

A2C (Advantage Actor-Critic) turns the value baseline from the previous project into a full algorithm: an actor network outputs the policy while a critic network learns V(s), and the critic's estimate is used to compute each action's advantage. Two ingredients make it work in practice — Generalized Advantage Estimation (GAE), which blends short-sighted one-step and full-return advantage estimates to trade off bias and variance, and running many copies of the environment in parallel so each gradient step sees a batch of decorrelated transitions instead of one highly-correlated trajectory. Running 8 parallel environments is enough to stabilize learning on LunarLander, a control task where the agent must fire thrusters to land a spacecraft gently between two flags. A2C is the synchronous sibling of A3C, which gathers experience from asynchronous workers instead of stepping them in lockstep.