The Kullback-Leibler (KL) divergence is a measure from information theory that quantifies how much one probability distribution differs from a second, reference distribution, expressed as the expected excess number of bits required to encode samples from the first using a code optimised for the second. It is non-negative, equal to zero only when the two distributions are identical, and is asymmetric, so it is a divergence rather than a true metric. KL divergence is central to machine learning, appearing in cross-entropy loss, variational inference, the evidence lower bound of variational autoencoders, and regularisation of policy updates in reinforcement learning.

Overview

  • KL divergence measures the information lost when one distribution is used to approximate another reference distribution.
  • It is defined as the expectation, under the first distribution, of the logarithmic ratio of the two distributions.
  • Because it is asymmetric and does not satisfy the triangle inequality, it is a divergence rather than a true distance metric.

Mechanisms

  • Non-negativity: KL divergence is always at least zero, equalling zero only when the distributions coincide (Gibbs’ inequality).
  • Asymmetry: the divergence from P to Q generally differs from Q to P, so direction matters in optimisation.
  • Decomposition: cross-entropy equals entropy plus KL divergence, linking it directly to common loss functions.
  • Estimation: it can be approximated from samples and is differentiable for gradient-based learning.

Applications

  • Cross-entropy loss for training classifiers.
  • The evidence lower bound regulariser in variational autoencoders.
  • Policy regularisation in reinforcement learning algorithms.
  • Knowledge distillation matching student and teacher distributions.

Provenance