Language model alignment is the set of techniques used to make a language model’s behaviour conform to human intentions, values and safety constraints. It typically follows pretraining with supervised fine-tuning and preference-based optimisation so that outputs are helpful, honest and harmless. Methods include reinforcement learning from human feedback and direct preference optimization.

Content

  • The dominant pipeline combines instruction tuning with preference optimisation derived from human or AI feedback, optionally augmented by constitutional rules and red-teaming. Alignment addresses both capability shaping (following instructions) and safety (refusing harmful requests, avoiding deception), and remains an active research frontier as models grow more capable.