Variational inference (VI) is a family of algorithms in Bayesian machine learning that approximates intractable posterior distributions p(z|x) by positing a simpler, tractable family of distributions q(z; φ) and optimising its parameters to minimise the Kullback-Leibler divergence from the true posterior, equivalently maximising the Evidence Lower BOund (ELBO) on the log marginal likelihood. By recasting probabilistic inference as an optimisation problem rather than a sampling problem, VI achieves scalability to large datasets and high-dimensional latent spaces that Markov Chain Monte Carlo methods cannot easily reach. The reparameterisation trick enables gradient-based ELBO optimisation through stochastic estimates, making VI the foundational inference engine for Variational Autoencoders, hierarchical generative models, and probabilistic programming systems. VI trades posterior exactness for computational tractability, and is the preferred method wherever fast, amortised, or online Bayesian inference is required.
Overview
- Why it exists: The central challenge in Bayesian machine learning is that posterior distributions p(z|x) are analytically intractable for virtually all models of practical interest — the normalising constant requires integrating over all latent states. Markov Chain Monte Carlo methods produce asymptotically exact samples but scale poorly to large datasets and high-dimensional spaces. Variational inference reframes inference as optimisation, unlocking Stochastic Gradient Descent and modern autodifferentiation frameworks.
- Core idea: Introduce a family of tractable distributions Q = {q(z; φ)} parameterised by variational parameters φ. Find the member of Q closest to the true posterior by solving: φ* = argmin KL(q(z; φ) || p(z|x)). Because the KL divergence involves the intractable posterior, VI instead maximises the ELBO: ELBO(φ) = E_{q}[log p(x,z)] − E_{q}[log q(z; φ)], which lower-bounds the log evidence log p(x).
- Historical trajectory: Early VI work (Hinton & van Camp 1993; Jaakkola & Jordan 1999) used coordinate-ascent under Mean-Field Approximation factorisation assumptions. Black-box VI (Ranganath et al. 2014) and the Reparameterisation Trick (Kingma & Welling 2013; Rezende et al. 2014) generalised VI to arbitrary differentiable models, enabling deep generative models. VI is now standard infrastructure in Probabilistic Programming frameworks such as Pyro, Stan (via ADVI), and TensorFlow Probability.
- Relationship to Expectation-Maximisation: EM can be viewed as a special case of VI where the approximate posterior is a point mass; VI generalises this to full distributional approximations with uncertainty quantification.
Key Mechanisms
- Evidence Lower Bound (ELBO)
- The primary optimisation target: ELBO(φ) = E_q[log p(x|z)] − KL(q(z;φ) || p(z))
- Decomposes into a reconstruction term and a regularisation term (KL penalty on the prior)
- Tightness: the ELBO equals log p(x) when q(z;φ) = p(z|x) exactly
- Maximising ELBO simultaneously improves fit to data and keeps q close to the prior
- Reparameterisation Trick
- Enables low-variance gradient estimation through q(z;φ) for continuous latent variables
- Rewrites z ~ q(z;φ) as z = g(ε, φ) where ε ~ p(ε) is a noise variable independent of φ
- Gradients ∇_φ ELBO can then be estimated via Monte Carlo: differentiating through sampled ε
- Critical for training Variational Autoencoder end-to-end via Backpropagation
- Mean-Field Approximation
- The most common variational family: q(z;φ) = ∏_i q_i(z_i; φ_i) — fully factored (independent) latent dimensions
- Tractable coordinate-ascent variational inference (CAVI) updates exist for exponential family models
- Known limitation: cannot capture posterior correlations between latent dimensions
- Normalising Flows
- Enrich the variational family beyond mean-field by composing invertible differentiable transformations
- Transform a simple base distribution (e.g., Gaussian) into a more expressive approximate posterior
- Variants: planar flows, radial flows, Real-NVP, Glow, Neural Spline Flows
- Enable posterior approximations that capture multi-modality and complex correlations
- Amortised Inference
- Rather than optimising φ separately for each datapoint, an inference network (encoder) learns a mapping x → φ(x) shared across data
- Dramatically accelerates inference at test time; underpins Variational Autoencoder architecture
- Trade-off: introduces amortisation gap between per-instance optimal φ and network predictions
- Black-Box Variational Inference (BBVI)
- Score-function (REINFORCE) estimator of ELBO gradients requires only samples from q and evaluations of the joint log p(x,z)
- Applicable to discrete and non-reparameterisable latent variables
- Higher variance than reparameterisation; mitigated via control variates, Rao-Blackwellisation
- Stochastic VI (SVI)
- Extends VI to large datasets via mini-batch subsampling of the ELBO’s data term
- Introduced by Hoffman et al. (2013) for Probabilistic Topic Modelling (Latent Dirichlet Allocation)
- Natural gradient updates in SVI converge faster than vanilla gradient descent for exponential family models
Applications and Use Cases
- Variational Autoencoder (VAE)
- Canonical deep generative model: encoder network approximates posterior q(z|x), decoder models p(x|z)
- Enables image generation, interpolation in latent space, semi-supervised learning, and anomaly detection
- Hierarchical extensions (NVAE, VDVAE, Ladder VAE) stack multiple latent levels for richer representations
- Probabilistic Topic Modelling
- Latent Dirichlet Allocation (LDA) and its successors use VI for scalable posterior inference over document-topic mixtures
- Stochastic VI enabled LDA to scale to millions of documents
- Bayesian Deep Learning
- Bayes by Backprop and similar methods place distributions over neural network weights using VI
- Enables calibrated uncertainty estimates in deep networks for safety-critical applications
- Approximate posteriors over weights allow predictive confidence intervals and active learning
- Probabilistic Programming
- Systems such as Pyro (PyTorch), NumPyro (JAX), Stan (ADVI), and TensorFlow Probability expose VI as a first-class inference backend
- Users specify generative models; VI automatically derives approximate posteriors
- Latent Diffusion Model
- Diffusion models trained in the latent space of a VAE encoder (e.g., Stable Diffusion) inherit the VI training objective
- VI pretraining of the encoder-decoder compresses images into a regularised, continuous latent space suitable for diffusion
- Natural language processing and Sequence Modelling
- VAE-based text models impose structure on sentence-level latent variables
- VI used in neural machine translation models for learning latent alignment distributions
- Molecular design and drug discovery
- Graph VAEs using VI learn latent molecular representations for generative chemistry
- Bayesian optimisation in latent space guided by VI-trained models accelerates lead compound identification
- Recommender systems
- Collaborative Variational Autoencoder and Variational Matrix Factorisation apply VI to implicit feedback data
- Uncertainty estimates from VI enable exploration-exploitation trade-offs in recommendations
- Reinforcement Learning
- Variational inference used in model-based RL (DREAMER, PlaNet) to infer compact world-model latent states from observations
- Enables planning and policy learning entirely in latent space
Standards and Context
- Frameworks: Pyro, NumPyro, TensorFlow Probability, Turing.jl, Stan (ADVI mode), Blackjax all provide production-grade VI implementations
- ELBO variants: Evidence Lower BOund (ELBO), Importance-Weighted ELBO (IWAE), Renyi-alpha divergence bounds, and the Free Energy are related objectives studied in the VI literature
- Evaluation: Held-out log-likelihood (estimated via importance-weighted sampling), FID for image generation, perplexity for topic models
- Key theoretical connections: VI is the free-energy minimisation principle from Statistical Physics; the ELBO corresponds to the negative variational free energy. This connection links VI to the Free Energy Principle in neuroscience (Friston).
- Convergence and guarantees: VI objective is generally non-convex; local optima are common. Mean-field CAVI has convergence guarantees for exponential family models but may converge to poor local optima for multi-modal posteriors.
- Active research directions: Tighter ELBO bounds (IWAE, VIMCO), implicit variational posteriors, adversarial VI (combining with GANs), discrete variational methods (Gumbel-Softmax, REINFORCE), and combining VI with Markov Chain Monte Carlo (e.g., MCMC-corrected VI, diffusion-based samplers).