Preference data is a dataset of paired or ranked comparisons in which human annotators indicate which of two or more model outputs they prefer, rather than providing an absolute quality score. It is the primary training signal for reward models used in reinforcement learning from human feedback, since relative judgements are typically easier and more consistent for annotators to produce than calibrated absolute ratings. The quality and diversity of preference data materially shape the behaviour that RLHF instils in a model.

Content

  • Preference data is a dataset of paired or ranked comparisons in which human annotators indicate which of two or more model outputs they prefer, rather than providing an absolute quality score. It is the primary training signal for reward models used in reinforcement learning from human feedback, since relative judgements are typically easier and more consistent for annotators to produce than calibrated absolute ratings. The quality and diversity of preference data materially shape the behaviour that RLHF instils in a model.