Length-Bias Audit
Key Insight
Length bias is the well-documented tendency of RLHF-tuned models to grow more verbose over training — not because longer answers are genuinely better, but because reward models and human preference data quietly correlate length with quality, and the policy learns to exploit that correlation. This project plots the completion-length distributions of your PPO- and DPO-trained models before and after tuning to make the drift visible. Why it matters: length bias is a concrete, easy-to-measure instance of reward hacking — the model is maximizing a flawed proxy rather than true helpfulness — and spotting it is the first step toward the length-corrected losses (such as SimPO) that try to remove it.