Skip to content
Generative AI

What does RLHF change about a base model?

Behaviour, not knowledge. Supervised fine-tuning on demonstrations, then a reward model trained on human comparisons of response pairs, then policy optimisation against it with a KL penalty to stay near the base. DPO skips the reward model.

datasciencetrivia.com

Card 1 of 213. Answer: Behaviour, not knowledge. Supervised fine-tuning on demonstrations, then a reward model trained on human comparisons of response pairs, then policy optimisation against it with a KL penalty to stay near the base. DPO skips the reward model.

Created by santiviquez

About · To suggest new questions or report an error send me a dm.

All questions