Skip to main content

SFT a Small Base Model

Key Insight​

Supervised fine-tuning (SFT) is the first post-training stage and the foundation every later RLHF step builds on: you take a base model — a raw next-word predictor fresh out of pretraining that only knows how to continue text — and fine-tune it on a dataset of (instruction, response) demonstrations so it learns to follow requests instead of merely autocompleting them (instruction tuning). Mechanically it is plain supervised learning: the demonstrated response is the answer key, and the model is trained with a cross-entropy loss to reproduce those tokens. This project fine-tunes a small base model (e.g. Qwen-0.5B) on a compact instruction dataset so you can watch a brilliant-but-aimless autocomplete turn into a usable assistant. Why it matters: SFT alone produces a decent model, and the resulting checkpoint becomes the frozen reference model that the reward model, PPO, and DPO stages all measure their changes against.