Generative AI
What does RLHF change about a base model?
Behaviour, not knowledge. Supervised fine-tuning on demonstrations, then a reward model trained on human comparisons of response pairs, then policy optimisation against it with a KL penalty to stay near the base. DPO skips the reward model.