Skip to main content

Dataset-Quality Study

Key Insight​

The same offline RL algorithm can look brilliant or useless depending only on who collected the data, and this study makes that dependence visible by running one fixed method across three D4RL datasets of the same task — random (a flailing behavior policy), medium (a half-trained one), and expert (a polished one) — and plotting final return against data quality. The lesson is that offline RL cannot conjure skill the data never shows: when the training data is already collected by an expert, a simple method like copying the expert's actions (behavior cloning) works incredibly well. However, when the data is poor or random, a true offline RL algorithm stands out because it can look at many mediocre attempts, find the best parts of each, and combine (or "stitch") them together into a single, high-performing strategy. This ability to create a policy that is better than any single trajectory in the dataset is what makes offline RL much more powerful than simple imitation. Knowing where your dataset sits on this curve tells you whether a sophisticated method is even worth the trouble.