Skip to main content

DPO

Key Insight​

Direct Preference Optimization (DPO) collapses the entire PPO-style RLHF pipeline into a single supervised loss on (chosen, rejected) pairs — no separate reward model, no rollouts, no RL loop. A clever mathematical shortcut (a closed-form derivation) proves that the model's own answer probabilities—how likely it is to output each word compared to a frozen reference model—already contain an implicit reward. This means we don't need a separate reward model at all: one simple training step just makes the human-preferred answer more likely and the rejected one less likely, while using the reference model comparison as a leash to keep the model from drifting into nonsense. This project trains DPO on the same preference data you used for PPO-RLHF and compares quality, training time, and stability. Why it matters: DPO is far simpler and cheaper to run than PPO, which is why it became the default in many open-source post-training pipelines — though its DPO-family variants (KTO, IPO, ORPO, SimPO) exist precisely because the plain loss has failure modes such as length bias.