Reward modelling is a machine-learning technique that trains a separate model to predict human preferences and use its scores as a reward signal for optimising another agent. It underpins reinforcement learning from human feedback, where the reward model ranks candidate outputs so a policy can be tuned toward preferred behaviour. It is central to aligning large language models with human values and intent.
Content
- A reward model is trained on human comparisons between outputs, then provides the reward in reinforcement learning from human feedback. Its quality bounds alignment: reward misspecification or hacking can cause models to optimise the proxy rather than the intended objective.