The Kullback-Leibler (KL) divergence is a measure from information theory that quantifies how much one probability distribution differs from a second, reference distribution, expressed as the expected excess number of bits required to encode samples from the first using a code optimised for the second. It is non-negative, equal to zero only when the two distributions are identical, and is asymmetric, so it is a divergence rather than a true metric. KL divergence is central to machine learning, appearing in cross-entropy loss, variational inference, the evidence lower bound of variational autoencoders, and regularisation of policy updates in reinforcement learning.
Overview
- KL divergence measures the information lost when one distribution is used to approximate another reference distribution.
- It is defined as the expectation, under the first distribution, of the logarithmic ratio of the two distributions.
- Because it is asymmetric and does not satisfy the triangle inequality, it is a divergence rather than a true distance metric.
Mechanisms
- Non-negativity: KL divergence is always at least zero, equalling zero only when the distributions coincide (Gibbs’ inequality).
- Asymmetry: the divergence from P to Q generally differs from Q to P, so direction matters in optimisation.
- Decomposition: cross-entropy equals entropy plus KL divergence, linking it directly to common loss functions.
- Estimation: it can be approximated from samples and is differentiable for gradient-based learning.
Applications
- Cross-entropy loss for training classifiers.
- The evidence lower bound regulariser in variational autoencoders.
- Policy regularisation in reinforcement learning algorithms.
- Knowledge distillation matching student and teacher distributions.