RLVR on Math
Key Insight
RL with Verifiable Rewards (RLVR) throws out the learned reward model entirely: when an answer can be checked by a program — a math result that matches the known solution, code that passes its unit tests — a deterministic verifier hands back an exact, unhackable reward for free. This project trains a small reasoning loop on a verifiable math subset (GSM8K-style) and watches the model's chain of thought grow longer over training, because writing out more reasoning steps leads to more correct — and therefore more rewarded — answers. Why it matters: RLVR is the engine of the reasoning-model wave (o1, R1), and it explains the grand bet of 2026 — the more domains we can convert into verifiable form, the more of RLHF's messy preference machinery we can throw away.