Skip to main content

PPO-Style RLHF

Key Insight​

This is the original RLHF recipe behind InstructGPT and ChatGPT: treat the prompt as the state, the generated completion as a sequence of actions, the reward model's score as a terminal reward, and run PPO to nudge the policy (the language model) toward higher-scoring completions. The catch is a KL-divergence penalty back to the frozen reference model: without that leash the policy quickly reward-hacks — producing gibberish the reward model happens to rate highly — so the penalty keeps it improving slowly toward what humans actually want. This project runs a mini-RLHF loop on a small model and small reward model while tracking the KL to the reference. Why it matters: PPO-RLHF is powerful but notoriously fiddly (a separate value head, advantage estimates on token-level returns, KL scheduling), and feeling that fragility firsthand is exactly why later methods like DPO and GRPO exist to remove moving parts.