Skip to main content

Decision Transformer

Key Insight​

The Decision Transformer throws out value functions and Bellman backups entirely and reframes offline RL as sequence modeling: feed a transformer a stream of (desired return-to-go, state, action) tokens and train it, exactly like a language model, to predict the next action. At test time you simply tell it the return you want — "from here, collect 500 reward" — and it autoregressively produces actions consistent with having achieved that, because during training it saw which action sequences led to which returns. This turns "find the optimal policy" into "predict what an agent that earned this much reward would do," which shines when the dataset is large and varied. Compared with IQL it needs no out-of-distribution machinery at all — the price is that it reliably hits only return targets the data actually demonstrates. Its value-free cousin that also models states and rewards is the Trajectory Transformer.