Implement IQL
Key Insight
IQL (Implicit Q-Learning) sidesteps the out-of-distribution problem more cleanly than CQL: instead of penalizing bad actions, it simply never asks Q about any action outside the dataset. It learns a value function V(s) with expectile regression — an asymmetric loss that leans toward the better outcomes the data already contains, approximating "the value of the best in-dataset action" without ever evaluating an unseen one — then trains the policy by advantage-weighted regression, copying dataset actions but weighting the good ones more heavily. Because every quantity is computed only on actions that really occurred, there is nothing to hallucinate, which makes IQL simpler and more robust than CQL and the modern default for offline RL.