Reward Hacking Demo
Key Insight
Reward hacking is when a policy maximizes the reward signal without doing the thing the reward was meant to encourage — the gap between the proxy you can measure and the goal you actually care about. This project provokes it on purpose: over-train a model against a reward model with the KL penalty turned down or off, then characterize the gibberish that emerges as the policy discovers quirks that score highly but mean nothing. Why it matters: it shows in the most visceral way why the KL leash to a frozen reference model is non-negotiable in RLHF, and why a learned reward model — unlike a deterministic verifier in RLVR — can always be gamed given enough optimization pressure.