Normalising flows are a class of generative models that learn complex probability distributions by composing a series of invertible, differentiable transformations that map a simple base distribution (typically Gaussian) to the target distribution, with exact log-likelihood computation via the change-of-variables formula and the Jacobian determinant. Both sampling and density evaluation are tractable.

Content

  • The theoretical foundations of normalising flows rest on the change-of-variables formula from probability theory, recognised in statistics for decades but operationalised for deep learning by Rezende and Mohamed (2015) and Dinh et al.’s NICE (2014) and RealNVP (2016) architectures. Early flows used simple coupling layers that partitioned the input and applied element-wise affine transforms, keeping the Jacobian triangular and thus its determinant O(d). This work established the key design constraint: the transformations must be invertible and their Jacobians must be efficiently computable.
  • Modern flow architectures include autoregressive flows (MAF, IAF) where each output dimension conditions on all previous ones — yielding expressive density estimators but slower sampling or evaluation in one direction; continuous normalising flows (CNF/FFJORD) that parameterise the transformation as an ODE and compute the log-determinant via Hutchinson’s trace estimator; and Glow, which extends RealNVP with invertible 1x1 convolutions for image generation. The key property distinguishing flows from VAEs and GANs is exact, tractable likelihood: the loss function is the average negative log-likelihood under the flow, computed without approximation.
  • Normalising flows matter for applications requiring reliable density estimates, not just samples. In scientific computing they serve as fast emulators for Bayesian posteriors, replacing expensive MCMC; in particle physics they accelerate simulation of detector responses; in finance they model joint return distributions with correct tail behaviour. For generative tasks, flows pioneered high-fidelity speech synthesis (WaveGlow) and musical audio (Glow-TTS), and they remain competitive in small-to-medium data regimes where the exact likelihood is more useful than raw sample quality.
  • In 2024–2025, normalising flows have been partially eclipsed in image generation by diffusion models but have found a robust niche in scientific machine learning and probabilistic programming. Neural spline flows, using piecewise-rational-quadratic transformations, have become a standard reference architecture. Flow matching (Lipman et al., 2022) and consistency models have emerged as efficient bridges between flows and diffusion, training continuous flows without ODE simulation during training. Integration with JAX (Distrax, Flowjax) enables JIT-compiled, GPU-vectorised inference over thousands of posterior samples simultaneously.