Sampling is the family of computational and statistical procedures that draw representative or informative values from probability distributions — ranging from classical Markov chain Monte Carlo mods (Metropolis-Hastings, Gibbs, Hamiltonian Monte Carlo, NUTS) and sequential mods (Sequential Monte…

Semantic Classification

Content

Compositional Relationships (Components)

SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:hasPart ai:MarkovChainMonteCarlo))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:hasPart ai:ImportanceSampling))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:hasPart ai:RejectionSampling))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:hasPart ai:SequentialMonteCarlo))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:hasPart ai:DiffusionSampling))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:hasPart ai:AutoregressiveSampling))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:hasPart ai:QuasiMonteCarlo))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:hasPart ai:NestedSampling))

## Dependency Relationships
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:requires ai:ProbabilityDistribution))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:requires ai:RandomNumberGenerator))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:requires ai:TargetDensity))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:requires ai:ProposalDistribution))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:requires ai:ConvergenceCriterion))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:dependsOn ai:InformationTheory))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:dependsOn ai:MeasureTheory))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:dependsOn ai:StochasticDifferentialEquations))

## Capability Relationships
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:enables ai:BayesianPosteriorInference))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:enables ai:GenerativeImageSynthesis))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:enables ai:LanguageModelDecoding))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:enables ai:UncertaintyQuantification))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:enables ai:SimulationBasedInference))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:enables ai:ModelSelection))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:supports ai:LargeLanguageModels))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:supports ai:DiffusionModels))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:supports ai:VariationalInference))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:supports ai:ParticleFilter))

## Implementation Relationships
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:implements ai:MetropolisHastingsAlgorithm))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:implements ai:GibbsSampling))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:implements ai:HamiltonianMonteCarlo))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:implements ai:NUTS))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:implements ai:DDIM))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:implements ai:DPMSolver))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:implements ai:NucleusSampling))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:implements ai:SpeculativeDecoding))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:uses ai:MarkovChains))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:uses ai:GradientInformation))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:uses ai:Entropy))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:uses ai:ScoreFunction))

## Reduction Relationships
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:reduces ai:SamplingVariance))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:reduces ai:AutocorrelationTime))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:reduces ai:InferenceCost))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:reduces ai:DecodingLatency))
SubClassOf(ai:Sampling
  ObjectSomeValuesFrom(ai:reduces ai:HallucinationRate))

## Annotations
AnnotationAssertion(rdfs:label ai:Sampling "Sampling"@en)
AnnotationAssertion(rdfs:comment ai:Sampling "Family of computational procedures drawing values from probability distributions spanning classical MCMC (Metropolis-Hastings, Gibbs, HMC, NUTS), sequential Monte Carlo, importance/rejection sampling, quasi-Monte Carlo (Sobol), nested sampling, autoregressive LLM decoding strategies (top-k, nucleus/top-p, temperature, min-p, mirostat, eta), and diffusion model schedulers (DDIM, DPM-Solver, Euler, Heun, Karras EDM), unifying all regimes under approximation of intractable expectations E_p[f(x)] via strategically chosen sample points."@en)
AnnotationAssertion(dcterms:identifier ai:Sampling "AI-2047"^^xsd:string)
AnnotationAssertion(dcterms:subject ai:Sampling "Monte Carlo Methods, Bayesian Inference, Language Model Decoding, Diffusion Samplers, Statistical Computing"@en)
DisjointClasses(ai:Sampling ai:DeterministicIntegration)
DisjointClasses(ai:Sampling ai:GridApproximation)

## Property Characteristics
AsymmetricObjectProperty(ai:requires)
AsymmetricObjectProperty(ai:enables)
AsymmetricObjectProperty(ai:implements)
TransitiveObjectProperty(ai:dependsOn)
FunctionalDataProperty(ai:varianceReductionFactor)

About Sampling

  • Sampling occupies a central position in modern artificial intelligence, statistics, and scientific computing.
  • Every practical inference task involving a probability distribution too complex for analytical treatment — posterior distributions in Bayesian modelling, data distributions learned by generative networks, the token distribution at each decoding step of an LLM, the reverse diffusion trajectory of a latent-space image model — reduces to a sampling problem: drawing realisations from p(x) or estimating integrals ∫ f(x) p(x) dx without evaluating every point in the (often infinite-dimensional) support of p.
  • The canonical approximation is the Monte Carlo estimate: given N independent draws {x₁, …, xₙ} ~ p(x), the expectation E_p[f(x)] ≈ (1/N) Σᵢ f(xᵢ) converges by the law of large numbers with standard error O(N^{-1/2}) regardless of dimensionality.
  • This is a dramatic improvement over deterministic quadrature requiring O(N^{d/m}) points for d-dimensional integrands of smoothness m — the curse of dimensionality.
  • The art of sampling is therefore variance reduction: designing procedures that produce draws whose empirical average converges to the true expectation faster than naive random sampling.
  • The fundamental challenge: the curse of dimensionality means that most probability mass concentrates in a thin “typical set” — not at the mode and not at the tails — making naive grid search or uniform sampling exponentially inefficient as d increases.
  • Ergodic theory underpins MCMC: the Birkhoff ergodic theorem guarantees that time averages (sample averages along a Markov chain trajectory) equal space averages (true posterior expectations) for ergodic chains, justifying asymptotic consistency of all MCMC estimators.
  • Central limit theorem for MCMC: under weak mixing conditions, (1/N) Σᵢ f(xᵢ) satisfies a CLT with asymptotic variance σ²_MCMC = σ²_iid × (1 + 2 Σ_k ρ_k), where ρ_k is the lag-k autocorrelation; the integrated autocorrelation time τ_int = 1 + 2 Σ_k ρ_k quantifies the effective sample size loss due to dependence.
  • Three major variance reduction paradigms: (1) improved proposals (HMC, NUTS — lower autocorrelation per sample); (2) control variates (Rao-Blackwellisation, ZV-MCMC — reducing variance of the estimator for fixed samples); (3) quasi-randomness (QMC, LHS — replacing random draws with deterministic low-discrepancy sequences).

Components/Architecture

Markov Chain Monte Carlo (MCMC)

  • MCMC constructs a Markov chain whose stationary distribution equals the target p(x) ∝ π(x), producing correlated samples that — after a burn-in period — constitute an ergodic approximation of the posterior.
  • The three foundational algorithms are Metropolis-Hastings, Gibbs, and Hamiltonian Monte Carlo.
  • Key theoretical property: ergodicity guarantees that time averages equal space (ensemble) averages, so correlated MCMC samples still produce consistent estimators as N → ∞.
  • Practical diagnostics: R-hat (Gelman-Rubin) convergence statistic comparing within-chain vs between-chain variance; effective sample size (ESS) = N / (1 + 2 Σ_k ρ_k) accounting for autocorrelation ρ_k at lag k; trace plots for stationarity; divergent transitions (HMC-specific, indicating numerical instability near posterior curvature).

Metropolis-Hastings (MH)

  • Proposed by Metropolis et al. (1953) for physics simulations; generalised by Hastings (1970, Cambridge).
  • Algorithm: propose x’ ~ q(x’|xₜ); accept with probability α = min(1, [π(x’) q(xₜ|x’)] / [π(xₜ) q(x’|xₜ)]); set xₜ₊₁ = x’ if accepted, else xₜ₊₁ = xₜ.
  • The acceptance ratio cancels the unknown normalisation constant Z (π(x) = p(x)/Z), making MH applicable whenever the unnormalised density π(x) is evaluable.
  • The random-walk proposal q(x’|xₜ) = N(x’; xₜ, σ²I) yields Random Walk Metropolis (RWM); optimal σ² scales as 2.38²/d (Roberts et al. 1997) giving acceptance rate ≈ 23.4% and mixing time O(d).
  • The independence Metropolis-Hastings variant uses a fixed proposal q(x’) independent of xₜ — efficient when q approximates p closely but prone to sticky behaviour when tails disagree.
  • Limitation: in d > 50 dimensions, random-walk steps become pathologically small, requiring exponentially many evaluations to traverse the posterior.
  • Adaptive Metropolis (Haario et al. 2001) learns the proposal covariance online from the chain history, converging toward the optimal covariance 2.38²/d × Σ_posterior.
  • Delayed rejection adaptive Metropolis (DRAM, Haario et al. 2006) combines adaptive proposals with a second proposal attempt on rejection, improving efficiency in complex posteriors.
  • Key correctness requirement: the proposal must be ergodic (capable of reaching any region of positive posterior mass) and reversible (detailed balance must hold); adaptive MH must use diminishing adaptation to maintain ergodicity.

Gibbs Sampling

  • Gibbs sampling (Geman & Geman 1984; popularised in statistics by Gelfand & Smith 1990) exploits factored conditionals: cycle over coordinates i = 1,…,d, drawing xᵢ ~ p(xᵢ | x_{-i}), which automatically satisfies detailed balance with acceptance probability 1.
  • When all full conditionals are available in closed form — as in conjugate Bayesian models (Normal-Normal, Gamma-Poisson, Dirichlet-Multinomial) — Gibbs is highly efficient and forms the backbone of BUGS/JAGS-style probabilistic programming.
  • Collapsed Gibbs (Rao-Blackwellisation) marginalises out subsets of variables analytically before sampling the rest, reducing variance and improving mixing.
  • Blocked Gibbs groups correlated variables into blocks and samples each block jointly, improving mixing vs coordinate-wise updates.
  • Limitations: slow mixing under strong posterior correlations (requiring many steps to traverse elongated posterior ridges); inapplicable when conditionals lack closed forms.
  • Systematic vs random-scan Gibbs: systematic scan (fixed coordinate order 1, 2, …, d) satisfies detailed balance only in pairs of scans; random-scan (randomly choose coordinate at each step) satisfies detailed balance exactly per step but has higher variance.
  • Auxiliary variable Gibbs: intractable conditionals can be made tractable by introducing auxiliary variables — the Albert-Chib probit model (augment with latent normals), slice sampling (augment with a uniform variable), and the Polya-Gamma trick for logistic regression (Polson et al. 2013).
  • Latent Dirichlet Allocation (LDA) collapsed Gibbs sampler (Griffiths & Steyvers 2004) is the classical topic modelling inference algorithm, marginalising word and document-topic distributions analytically and sampling topic assignments for each word in O(K) per token — the default algorithm in Mallet and Gensim.
  • Restricted Boltzmann Machines (RBMs) use block Gibbs between visible and hidden units; contrastive divergence (Hinton 2002) approximates the gradient of log-likelihood using one or a few Gibbs steps, enabling practical training of energy-based models.

Hamiltonian Monte Carlo (HMC)

  • HMC (Duane et al. 1987; Neal 2011) introduces auxiliary momentum variables p ~ N(0, M) and simulates Hamiltonian dynamics H(x, p) = -log π(x) + (1/2) pᵀ M⁻¹ p on the augmented space (x, p).
  • Uses L leapfrog steps of step-size ε to propose a distant, low-autocorrelation point; a final MH accept/reject corrects for numerical discretisation error.
  • The leapfrog integrator is volume-preserving and time-reversible, ensuring detailed balance.
  • HMC exploits gradient information ∇_x log π(x), dramatically reducing random-walk behaviour: mixing time scales as O(d^{1/4}) vs O(d) for RWM.
  • The mass matrix M (preconditioning) and choice of (L, ε) critically affect efficiency; hand-tuning is replaced in practice by the No-U-Turn Sampler.
  • Riemannian manifold HMC (Girolami & Calderhead 2011, JRSS-B) adapts M using the Fisher information metric G(x) at each position, dramatically improving mixing on highly correlated posteriors — at the cost of Christoffel symbol computations O(d³) per leapfrog step.

No-U-Turn Sampler (NUTS)

  • NUTS (Hoffman & Gelman 2014) eliminates the need to specify L by dynamically building a balanced binary tree of leapfrog trajectory points, stopping when the trajectory “turns back” on itself (∂H/∂t < 0, a U-turn).
  • A slice variable enforces detailed balance across the tree’s leaves.
  • NUTS also adapts ε via dual-averaging (Nesterov 2009) during warm-up.
  • The result is a parameter-free, self-tuning HMC variant achieving O(d^{1/4}) mixing with no user effort.
  • NUTS is the algorithm underpinning Stan, PyMC, NumPyro, and TensorFlow Probability.
  • Empirical studies show NUTS requires 2-10× fewer gradient evaluations than well-tuned HMC and 100-1000× fewer than RWM on problems with d > 100.
  • Divergent transitions in NUTS signal that the leapfrog trajectory entered a region of extreme posterior curvature; they are a critical diagnostic indicating model mis-specification, inadequate parameterisation, or step-size too large — not merely a tuning issue.
  • Non-centred parameterisations (e.g., replacing θ ~ N(μ, σ²) with θ = μ + σ·η where η ~ N(0,1)) dramatically reduce funnel-shaped posteriors common in hierarchical models, enabling NUTS to traverse the full posterior efficiently.
  • NUTS in practice: Stan’s default of 1000 warm-up + 1000 sampling iterations is usually sufficient; R-hat < 1.01 and ESS > 100 per parameter are standard convergence criteria per Vehtari et al. (2021).
  • MCLMC (Robnik & Seljak 2023, Microcanonical Langevin Monte Carlo): a recent improvement replacing the momentum resampling step with a continuous Langevin dynamics conserving the microcanonical ensemble, achieving up to 10× speedup over NUTS on high-dimensional posteriors; implemented in BlackJAX.

Importance Sampling (IS) and Self-Normalised Variants

  • Importance sampling estimates E_p[f(x)] by drawing from a tractable proposal q(x) and reweighting: E_p[f(x)] ≈ Σᵢ wᵢ f(xᵢ) / Σᵢ wᵢ where wᵢ = p(xᵢ)/q(xᵢ) and {xᵢ} ~ q(x).
  • Unbiased when q(x) > 0 wherever p(x) > 0 (absolute continuity).
  • Variance depends critically on the effective sample size ESS = (Σ wᵢ)² / Σ wᵢ² ∈ [1, N]; IS degrades catastrophically in high dimensions when q has lighter tails than p, producing weight collapse.
  • Self-normalised IS (SNIS) cancels the unknown normalisation Z via the ratio of weighted sums, introducing bias O(1/N) but often lower MSE in practice.
  • Annealed IS (Neal 1998) bridges between simple q and complex p through a geometric tempering path p_t ∝ p^t q^{1-t}, t ∈ {0, 1/(T-1), …, 1}, accumulating weights multiplicatively — a key tool for marginal likelihood estimation.
  • Pareto-smoothed IS (Vehtari et al. 2019, PSIS-LOO) fits a Pareto distribution to the tail of weights to diagnose and correct weight instability; the Pareto shape parameter k̂ > 0.7 signals unreliable IS estimates and is a standard diagnostic in Stan and LOO cross-validation.

Rejection Sampling

  • Rejection sampling draws x ~ q(x) and accepts with probability p(x) / (M q(x)) where M ≥ sup_x p(x)/q(x) is a bounding constant, yielding exact samples from p(x).
  • Acceptance rate is 1/M; finding a tight M in high dimensions is difficult as M typically grows exponentially with d, making rejection sampling effective only in d ≤ 10-20.
  • Adaptive rejection sampling (Gilks & Wild 1992) constructs piecewise-exponential upper envelopes for log-concave targets.
  • Zig-Zag and bouncy particle samplers (Bouchard-Côté et al. 2018) extend rejection ideas to continuous-time piecewise-deterministic Markov processes that scale to higher dimensions.
  • Key application: sampling from truncated distributions (e.g., truncated-Normal for constrained parameters) where the truncation region is small enough that the accept rate remains reasonable.
  • Slice sampling (Neal 2003) introduces a uniform auxiliary variable u ~ Uniform(0, π(xₜ)) and samples x’ uniformly from the “slice” {x : π(x) > u} — an implicit rejection sampler that automatically adapts its step size to local posterior curvature.
  • Elliptical slice sampling (Murray, Adams & MacKay 2010) handles Gaussian prior structure by proposing points on an ellipse defined by the current position and a Gaussian draw, enabling efficient sampling for GP-based models and Bayesian neural networks.

Sequential Monte Carlo (SMC)

  • SMC (Del Moral, Doucet & Jasra 2006; Chopin 2002) maintains a population of N weighted particles {(xᵢ, wᵢ)} propagated, reweighted, and resampled through a sequence of intermediate distributions π₁, …, πT bridging between a tractable prior π₁ and the target πT.
  • Steps per iteration: (1) Importance weight update wᵢ ∝ πₜ(xᵢ)/π_{t-1}(xᵢ); (2) Resample particles proportional to weights (multinomial, systematic, or stratified) when ESS < N/2; (3) MCMC rejuvenation via several MH/HMC steps to diversify the population.
  • SMC provides an unbiased estimate of the normalising constant Z = ∫ π(x) dx at each step — valuable for Bayesian model comparison via Bayes factors.
  • In SMC² (Chopin, Papaspiliopoulos & Rousset 2013), an outer SMC operates over parameter particles while an inner particle filter handles state-space latent variables.
  • SMC is the dominant method for real-time state estimation (robot navigation via the Particle Filter, financial filtering) and increasingly for Bayesian deep learning with Bayesian Neural Networks.
  • Tempering schedules: the sequence of bridging distributions can use geometric tempering (power posteriors), likelihood tempering (data tempering), or data-driven tempering — the choice strongly affects the efficiency of the normalising constant estimate.
  • Resample-move: systematic resampling (van der Merwe & Wan 2001) with stratified assignments minimises resampling variance; the number of distinct particles after resampling approximates ESS, so systematic resampling produces floor(ESS) distinct particles vs multinomial’s stochastic floor.
  • SMC for likelihood-free inference (ABC-SMC): replaces exact likelihood evaluation with a summary statistic distance threshold ε; iterates decreasing ε values, using SMC to adapt the population toward the target posterior p(θ | d(S(x_sim), S(x_obs)) < ε) — widely used in systems biology (Toni et al. 2009).
  • Parallel SMC: particle populations are embarrassingly parallelisable (each particle independently propagated); GPU implementations with N=10K-100K particles achieve real-time Bayesian inference for robotics and econometrics.

Nested Sampling

  • Nested sampling (Skilling 2004, 2006) targets the marginal likelihood Z = ∫ L(x) π(x) dx directly by transforming the d-dimensional integral into a 1D integral over the prior mass X(λ) = Pr(L(x) > λ).
  • Starting from N “live” particles drawn from π(x), it iteratively replaces the particle with lowest likelihood L_min with a new draw constrained to L > L_min, each step peeling off a thin shell of prior volume ΔX ≈ e^{-i/N}.
  • The evidence accumulates as Z = Σᵢ Lᵢ ΔXᵢ; uncertainty on Z is estimated via the variance over the dead points.
  • MultiNest (Feroz & Hobson 2009) uses ellipsoidal bounding for the constrained sampling step; PolyChord (Handley et al. 2015, Cambridge) uses slice sampling within the constrained likelihood region, scaling to d ~ 1000 dimensions; dynesty (Speagle 2020) implements dynamic nested sampling adaptively allocating particles.
  • Widely used in astrophysics (Planck CMB parameter estimation), cosmology, and molecular dynamics where log Z differences determine Bayes factors between physical models.

Latin Hypercube and Quasi-Monte Carlo Sampling

  • Latin Hypercube Sampling (LHS) (McKay, Beckman & Conover 1979) partitions each of d dimensions into N equal-probability strata and samples exactly once from each stratum per dimension, ensuring full marginal coverage.
  • Variance of LHS estimators is ≤ variance of pure Monte Carlo for any integrand with finite variance; improvements are greatest for monotone integrands.
  • Extensively used in engineering uncertainty quantification (UQ), surrogate model construction, and design of computer experiments for costly simulation codes.
  • Optimised LHS (Morris & Mitchell 1995) maximises the minimum pairwise distance between sample points (maximin criterion), improving space-filling beyond stratified random sampling; used in the Python pyDOE and SALib libraries.
  • Quasi-Monte Carlo (QMC) replaces pseudo-random draws with deterministic low-discrepancy sequences whose empirical distribution covers the unit hypercube more uniformly than random samples.
  • The Sobol sequence (Sobol 1967) uses a base-2 construction generating points with star-discrepancy D*_N = O((log N)^d / N), yielding integration error O((log N)^d / N) vs O(N^{-1/2}) for Monte Carlo — potentially O(N^{-1}) convergence for smooth integrands.
  • Halton sequences and digital nets provide alternative constructions; Randomised QMC (RQMC, Owen 1997) adds a random scramble preserving low-discrepancy whilst enabling variance estimation.
  • QMC is standard in financial option pricing, path-traced rendering, and sensitivity analysis (Sobol’ indices); modern applications include QMC-based stochastic gradient estimation (Buchholz et al. 2018) and quasi-random data augmentation.
  • Sobol’ sensitivity indices (Sobol 1993) decompose the variance of f(x) into contributions from individual variables and their interactions, enabling global sensitivity analysis; they are computed via IS-weighted QMC integrals and are the dominant method for uncertainty importance ranking in engineering simulation.
  • Dimension reduction for QMC: the effective dimension of an integrand determines whether QMC outperforms MC; the ANOVA-HDMR decomposition (Hoeffding 1948; Liu & Owen 2006) identifies low-effective-dimension structure enabling QMC to achieve near-O(N^{-1}) convergence even in formally high-d settings.
  • Deep learning applications of QMC: batch-normalised Sobol sequences improve training convergence in neural architecture search (NAS); QMC-based Bayesian hyperparameter optimisation outperforms random search in Gaussian process BO with d ≤ 20.

MCMC Diagnostics and Convergence Assessment

  • R-hat (Gelman-Rubin statistic): compares within-chain variance to between-chain variance across M parallel chains. R-hat = sqrt((N-1)/N + B/(N·W)) where B is between-chain and W is within-chain variance; R-hat < 1.01 is the 2021 standard (Vehtari et al.), superseding the older < 1.1 threshold.
  • Effective Sample Size (ESS): ESS = N / τ_int where τ_int = 1 + 2 Σ_k ρ_k is the integrated autocorrelation time; ESS_bulk measures mixing in the bulk of the distribution; ESS_tail (rank-normalised, Vehtari et al. 2021) specifically monitors tail mixing critical for posterior interval estimation.
  • Trace plots: visual inspection of the Markov chain trajectory over iterations; a stationary, well-mixing chain should resemble “hairy caterpillar” — dense oscillations around a stable mean with no visible trends or long-range correlations.
  • Autocorrelation function (ACF) plots: ACF(k) = Cov(f(xₜ), f(xₜ₊ₖ)) / Var(f(xₜ)); should decay quickly toward zero for well-mixing chains. Slow decay (ACF > 0.1 at lag k > 20) signals poor mixing.
  • Energy diagnostic (E-BFMI): Bayesian fraction of missing information from the energy — E-BFMI = Var(H(xₜ₊₁) - H(xₜ)) / Var(H(xₜ)) < 0.3 indicates poor HMC energy exploration; values > 0.3 are generally acceptable.
  • Posterior predictive checks (PPC): simulate data from the fitted posterior and compare its distribution to observed data; mismatches reveal model misspecification that may also compromise sampler convergence.
  • Geweke diagnostic: z-score comparing means of the first 10% vs last 50% of the chain; non-stationarity indicates failed convergence.
  • Heidelberger-Welch test: tests stationarity of each chain using a Cramer-von Mises statistic; useful for single long chains where R-hat is unavailable.

Variance Reduction and Control Variates

  • Control variates (CV): if a function g(x) with known mean E_p[g] is correlated with f(x), the estimator f̂_CV = (1/N) Σᵢ [f(xᵢ) - c(g(xᵢ) - E_p[g])] has lower variance for optimal c = Cov(f,g)/Var(g); widely used in MCMC via zero-variance (ZV) estimators with polynomial control variates.
  • Rao-Blackwellisation: replacing a Monte Carlo estimator by its conditional expectation E[f(x) | x_{-i}] (analytically computed) reduces variance without bias; the fundamental principle behind collapsed Gibbs and marginalised particle filters.
  • Antithetic variates: if x ~ p(x) is paired with a negatively correlated draw x̃ (e.g., x̃ = 2μ - x for symmetric p), the estimator (f(x) + f(x̃))/2 has variance ≤ Var(f(x))/2 when Cov(f(x), f(x̃)) < 0.
  • Stratified sampling: partition the support into K strata, draw Nₖ samples from each stratum with Σ Nₖ = N, weight by stratum probability; variance = Σₖ (pₖ²/Nₖ) σₖ² ≤ Var[MC] for any f with finite variance.
  • Coupling from the past (CFTP): (Propp & Wilson 1996) exact MCMC algorithm that provably produces draws from the exact stationary distribution (not asymptotically) by running chains backward from -∞; requires monotone or bounding chains; practically useful for finite discrete state spaces.

Probabilistic Programming Languages and Sampling Infrastructure

  • Stan (Columbia, 2012-present): probabilistic programming language with NUTS as the default sampler; HMC-ready automatic differentiation via its Math Library; interfaces in R (RStan), Python (PyStan), Julia (Stan.jl), and CmdStan; 200K+ users in statistics, ecology, political science, and clinical trials.
  • PyMC (v5, 2023): Python probabilistic programming with NUTS via Aesara/PyTensor autodiff; supports custom samplers (DEMetropolis, SMC); tight integration with ArviZ for convergence diagnostics and posterior visualisation.
  • NumPyro (Uber Research, 2019-present): NumPy-compatible probabilistic programming in JAX enabling JIT-compiled NUTS with hardware acceleration; NUTS on GPU/TPU achieves 10-100× speedup over CPU Stan for large hierarchical models (d > 1000).
  • BlackJAX (DeepMind/INRIA collaboration): modular JAX library of MCMC kernels (NUTS, SGLD, MALA, HMC, MCLMC, AIES); composable via JAX’s functional API; state-of-the-art MCLMC implementation achieving 10× speedup over NUTS on d > 500 posteriors.
  • Pyro / Numpyro (Uber AI, 2017-present): deep probabilistic programming integrating PyTorch/JAX neural networks with MCMC and SVI; supports structured variational inference, Gaussian process models, and normalising flow posteriors.
  • Turing.jl (University of Cambridge, Alan Turing Institute): Julia probabilistic programming with composable samplers (NUTS, HMC, MH, Gibbs, SMC, PG); native support for models with intractable likelihoods via ABC; 3K+ GitHub stars.
  • Infer.NET (Microsoft Research Cambridge): message-passing inference engine for factor graphs; supports exact and approximate inference (EP, VMP, Gibbs sampling) on discrete and continuous models; used for Xbox recommendation systems and clinical trial design.
  • ArviZ: Python library for Bayesian inference diagnostics and visualisation; computes R-hat, ESS_bulk, ESS_tail, MCSE, ELPD LOO, PSIS-LOO; the standard interface between Stan/PyMC/NumPyro outputs and post-sampling analysis.
  • Diffusers (Hugging Face): Python library implementing all major diffusion samplers (DDIM, DPM-Solver++, Euler, Heun, PNDM, UniPC, LMS) as interchangeable scheduler objects compatible with Stable Diffusion, DALL-E 3, Stable Diffusion XL, Stable Video Diffusion; 25K+ GitHub stars.
  • ComfyUI (Node-Based Diffusion Pipeline Interface): node-based diffusion model interface (30M+ users) exposing all major samplers (Euler, Euler_a, Heun, DPM-Solver++, LCM, DDIM) and schedulers (Karras, exponential, cosine, linear) as graph nodes enabling visual composition of sampling pipelines.

Stochastic Gradient and Scalable MCMC

  • Stochastic Gradient Langevin Dynamics (SGLD) (Welling & Teh 2011): combines SGD noise with Langevin dynamics — θₜ₊₁ = θₜ + (ε/2) ∇_θ log p(θ|x̃) + N(0, ε) where x̃ is a minibatch; converges to the posterior as ε→0 without an accept/reject step; scales to millions of data points.
  • SGHMC (Chen et al. 2014): adds momentum to SGLD, improving mixing; stochastic gradient noise acts as a thermostatic friction term, enabling principled Bayesian inference with minibatches.
  • Stein Variational Gradient Descent (SVGD) (Liu & Wang 2016): a deterministic particle method that transports a set of N particles to approximate the posterior by minimising KL divergence via functional gradient descent in an RKHS; bridges between MCMC and variational inference.
  • Scalable nested sampling: PolyChord’s GPU-accelerated likelihood evaluations and parallel live point propagation enable nested sampling on d ~ 1000 problems relevant to deep neural network weight-space posteriors (Cobb & Jalaian 2021).

Use Cases / Major Families

FamilyRepresentative AlgorithmPrimary DomainKey Scaling Property
MCMCNUTS (Stan/PyMC)Bayesian inferenceO(d^{1/4}) mixing
IS/SMCSequential MC²State filteringUnbiased log Z
NestedPolyChordCosmology/physicsEvidence estimation
QMCSobol + RQMCNumerical integrationO((log N)^d/N) error
LLM decodingNucleus + min-pText generationToken-adaptive thresholding
SpeculativeEAGLE-2LLM throughput2-4× speedup exact
Diffusion ODEDPM-Solver++(2M)Image synthesis10-20 NFE
Diffusion determ.DDIM η=0Image editing/inversionInvertible

Autoregressive LLM Sampling Strategies

  • Large language models parameterise a factored probability p(x₁, …, xT) = Π p(xₜ | x_{<t}) over discrete token vocabulary V (|V| ≈ 32,000-100,000).
  • At each step the model outputs a logit vector z ∈ ℝ^{|V|}, converted to probabilities via softmax; the decoding strategy determines how the next token is selected.
  • The choice of sampling strategy has profound effects on text quality, diversity, and coherence — pure greedy decoding (argmax) produces degenerate repetitive text; uniform random sampling produces incoherent output; practical strategies occupy the middle ground.

Temperature Scaling

  • Temperature T ∈ (0, ∞) divides logits before softmax: p_T(xₜ) = softmax(z/T).
  • T→0 collapses to greedy argmax (deterministic, high repetition); T=1 recovers the model distribution; T>1 flattens toward uniform (more diverse, less coherent).
  • Temperature is the foundational parameter and is composed with all other strategies.
  • Typical values: T=0.7-1.0 for creative generation, T=0.2-0.5 for factual/code tasks.
  • Temperature > 1.2 risks grammatical errors and factual hallucinations in most models.

Top-k Sampling

  • Top-k (Fan et al. 2018; “Hierarchical Neural Story Generation”, ACL 2018) restricts sampling to the k most probable tokens at each step, redistributing the filtered probability mass.
  • Addresses pathological samples from the tail of the distribution. k=50 is a common default; k=1 reduces to greedy.
  • Limitation: a fixed k is contextually inappropriate — when the distribution is peaked (one highly probable token), k=50 includes many unlikely tokens; when the distribution is flat, k=50 may be too restrictive.
  • Top-k was the dominant decoding strategy from 2018-2020 before being superseded by nucleus sampling.

Nucleus (Top-p) Sampling

  • Nucleus sampling (Holtzman et al. 2020, “The Curious Case of Neural Text Degeneration”, ICLR 2020) selects the smallest set Vₚ ⊆ V such that Σ_{v∈Vₚ} p(v|x_{<t}) ≥ p, then samples from the renormalised distribution over Vₚ.
  • p ∈ (0.9, 0.95) are standard values; typical defaults in OpenAI, Anthropic, and Mistral APIs.
  • Unlike top-k, nucleus adapts dynamically: when the model is confident (peaked distribution) only a few tokens are included; when uncertain (flat distribution) many tokens are included.
  • This closely tracks the model’s effective branching factor and produces more natural, less repetitive text than top-k.
  • Top-p is the default decoding strategy in most production LLM APIs and the Hugging Face transformers library; ICLR 2020 best paper candidate with 3000+ citations.

Min-p Sampling

  • Min-p sampling (Bradley et al. 2024, arXiv:2407.01082) sets a probability floor: exclude token v if p(v|x_{<t}) < min_p × max_{v’} p(v’|x_{<t}).
  • Only tokens whose probability is at least a fraction min_p of the most probable token’s probability are retained.
  • Unlike top-p which is reference-agnostic, min-p scales adaptively with model confidence.
  • When the model is confident (top token p=0.9), a min_p=0.05 threshold keeps tokens with p ≥ 0.045; when uncertain (top token p=0.05), only very improbable tokens are excluded.
  • Min-p outperforms top-p on creative generation benchmarks whilst avoiding the incoherence of high-temperature sampling.
  • Merged into llama.cpp and Hugging Face transformers in 2024; becoming the community-preferred strategy for creative writing tasks with instruction-tuned models.

Mirostat Adaptive Sampling

  • Mirostat (Basu et al. 2021, ICLR 2021) frames decoding as a feedback control problem targeting a fixed perplexity level μ.
  • At each step it estimates the text’s current perplexity and adjusts a dynamic top-k threshold k to drive perplexity toward μ.
  • Mirostat-2 simplifies this to a running estimate of the surprise Sₜ = -log p(xₜ|x_{<t}), adjusting a probability threshold τ to maintain E[S] ≈ μ.
  • The result is text whose information density is approximately constant — avoiding both repetitive low-surprise sequences and incoherent high-surprise token sequences.
  • Mirostat is particularly effective for extended generation (stories, essays) where top-p tends to degrade after hundreds of tokens due to compounding distributional drift.
  • Hyperparameter μ controls the target perplexity; μ=5 approximately replicates top-p=0.9 behaviour; μ=3 produces more focused text.

Eta Sampling

  • Eta sampling (Hewitt et al. 2022, EMNLP Findings 2022, arXiv:2210.15191) removes tokens below a threshold η = min(η_min, √(ε × e^{-H(p)})) where H(p) is the token distribution entropy and ε is a hyperparameter.
  • When the model is very confident (low H), the threshold tightens; when uncertain (high H), the threshold relaxes.
  • Eta sampling provides a principled information-theoretic motivation for adaptive truncation.
  • Particularly robust under distribution shift between training and inference — unlike top-p which assumes calibrated probability estimates.
  • Subsumes many properties of min-p and top-p under a unified entropy-based framework.

Speculative Decoding

  • Speculative decoding (Leviathan et al. 2023, ICML 2023; Chen et al. 2023, arXiv:2302.01318) addresses the autoregressive throughput bottleneck.
  • Each token requires a full forward pass through the large target model (typically 7B-70B parameters).
  • A small draft model (70M-1B parameters, 3-8× faster) autoregressively generates K candidate tokens.
  • The large model evaluates all K tokens in a single parallel forward pass and accepts/rejects each via a modified Metropolis-Hastings acceptance rule.
  • The acceptance rule guarantees the accepted sequence distribution matches the target model’s distribution exactly (not approximately) — zero quality loss.
  • Accepted tokens are appended; the first rejected token is resampled from the corrected distribution p_target - p_draft (renormalised).
  • Typical speedup: 2-4× wall-clock reduction; EAGLE-2 achieves 3.5× on Llama-3.1-70B.
  • Variants:
    • Medusa (Cai et al. 2024): adds K parallel decoding heads to the target model itself, avoiding a separate draft model.
    • EAGLE / EAGLE-2 (Li et al. 2024): feature-level draft model operating in the target model’s representation space.
    • SpecTr (Sun et al. 2023): tree-structured speculative drafts accepting multiple branches in parallel.
    • Lookahead decoding (Fu et al. 2024): Jacobi iteration to generate parallel draft tokens without a separate model.

Diffusion Model Sampling Strategies

  • Diffusion models (DDPM, Song et al. score matching, Rombach et al. LDM / Stable Diffusion Image Model) define a forward noising process q(xₜ|x₀) = N(xₜ; √ᾱₜ x₀, (1-ᾱₜ)I) that gradually adds Gaussian noise over T = 1000 steps.
  • A neural score network ε_θ(xₜ, t) ≈ -√(1-ᾱₜ) ∇_{xₜ} log p(xₜ) is learned to reverse the process.
  • Sampling = running the reverse process from xT ~ N(0, I) to x₀.
  • The schedule, step count, and ODE/SDE solver together determine both quality and speed; typical 2025 production trade-off is 20-30 NFE (network function evaluations) for high quality, 4-8 NFE for fast preview.

DDIM — Denoising Diffusion Implicit Models

  • DDIM (Song, Meng & Ermon 2020, ICLR 2021, arXiv:2010.02502) reinterprets DDPM’s Markov chain as a deterministic ODE.
  • Update rule: xₜ₋₁ = √ᾱₜ₋₁ (xₜ - √(1-ᾱₜ) ε_θ(xₜ,t)) / √ᾱₜ + √(1-ᾱₜ₋₁) ε_θ(xₜ,t).
  • The η parameter interpolates between fully deterministic (η=0, same latent → same image) and fully stochastic (η=1, recovers DDPM).
  • η=0 enables DDIM inversion: encode real images into latent space for editing; the inversion is approximate but sufficient for text-guided editing.
  • Supports 10-50 step generation vs DDPM’s 1000 steps with near-identical perceptual quality.
  • DDIM inversion is the backbone of SDEdit, prompt-to-prompt, null-text inversion, and most diffusion image editing pipelines.

DPM-Solver and DPM-Solver++

  • DPM-Solver (Lu et al. 2022a, NeurIPS 2022, arXiv:2206.00927) derives fast ODE solvers by expressing the diffusion ODE in the log-SNR domain λ(t) = log(ᾱₜ/(1-ᾱₜ)) and applying exponential integrators.
  • DPM-Solver-2 achieves high-quality samples in 10-20 steps (global truncation error O(h²)); DPM-Solver-3 in 10-15 steps.
  • DPM-Solver++ (Lu et al. 2022b, arXiv:2211.01095) adds multistep history terms; DPM-Solver++(2M) (M=2 multistep) is the most widely deployed fast sampler in Stable Diffusion WebUI (AUTOMATIC1111), Node-Based Diffusion Pipeline Interface, and InvokeAI.
  • Converges in the order-2 sense with global truncation error O(h²) vs DDIM’s O(h) (first-order).

Euler, Euler-Ancestral, and Heun Methods

  • Euler is the first-order explicit integrator for the diffusion ODE: xₜ₋₁ = xₜ + (tₜ₋₁ - tₜ) d/dt[xₜ]; fast (1 NFE/step) but requires many steps for high quality.
  • Euler-a (ancestral) adds Gaussian noise at each step scaled to the remaining noise schedule, recovering the DDPM stochastic reverse process; produces diverse outputs but non-invertible.
  • Heun’s method (second-order Runge-Kutta corrector) evaluates the score network twice per step: a predictor step (Euler) followed by a corrector using the averaged score, halving the global error at the cost of 2× NFE.
  • In Karras et al. (2022) notation, the Heun solver on the EDM ODE is the default high-quality reference sampler.

Karras EDM and Restart Sampling

  • Karras et al. (2022) (“Elucidating the Design Space of Diffusion-Based Generative Models”, NeurIPS 2022, arXiv:2206.00364) provide a unified analysis identifying the optimal noise schedule σ(t), preconditioning scalings, and ODE/SDE formulation.
  • The EDM framework separates the sampler design space into: noise schedule (σ_min, σ_max, ρ), ODE solver (Euler, Heun, RK45), and stochastic correction (Langevin noise injection).
  • Restart sampling (Xu et al. 2023, NeurIPS 2023) adds backward stochasticity by periodically jumping back to a higher noise level and re-running the deterministic ODE — trading extra NFE for diversity and reduced mode collapse.
  • The Karras et al. analysis is the theoretical foundation of ComfyUI’s sampling pipeline and informs most 2023-2026 diffusion sampler research.
  • The EDM preconditioning (c_skip, c_out, c_in, c_noise scalings) eliminates the need for cosine or linear noise schedules, providing near-optimal variance throughout the trajectory.

Sampling vs Variational Inference: Trade-offs

  • Variational inference (VI) approximates the posterior by optimising a parametric family q_φ(x) to minimise KL(q_φ || p) via gradient descent — deterministic, fast, but biased (mean-field VI underestimates posterior variance).
  • MCMC produces asymptotically exact samples but is slow for large models; the choice between VI and MCMC depends on whether speed or accuracy is the binding constraint.
  • Normalising flows as posterior approximators bridge the gap: flows are expressive enough to match complex posteriors while remaining amenable to reparameterisation-trick gradient estimation — Variational Inference with Normalising Flows (Rezende & Mohamed 2015) is the dominant approach in modern probabilistic programming.
  • Pathfinder (Zhang et al. 2022, Stan ecosystem): rapid VI algorithm that initialises from the L-BFGS optimisation trajectory, providing a warm start for NUTS and reducing warm-up by 10-100× on complex posteriors.
  • MCMC with VI initialisation: using a VI approximation as a proposal for IS (importance weighted autoencoder, IWAE) or as a warm start for MCMC chains (Stan’s default from 2023) combines the efficiency of VI with the asymptotic correctness of MCMC.
  • Hybrid approaches (MCMC-corrected VI, VBMC — Acerbi 2018): use VI to identify high-probability regions, then MCMC to refine; VBMC specifically targets expensive likelihoods with O(100) evaluations (e.g., neuroscience, psychology psychometric models).

Sampling in the Diffusion Model Ecosystem: Full Pipeline

  • Training: diffusion models are trained via denoising score matching — minimise E[||ε_θ(√ᾱₜ x₀ + √(1-ᾱₜ) ε, t) - ε||²] where ε ~ N(0,I) and t ~ Uniform(1, T); the noise schedule {ᾱₜ} (cosine, linear, EDM σ(t)) determines the difficulty of denoising at each step.
  • Guidance mechanisms: classifier guidance (Dhariwal & Nichol 2021) adds ∇_x log p(y|x) to the score, steering samples toward class y at inference; classifier-free guidance (Ho & Salimans 2022) trains a joint model with randomly dropped conditioning c, using CFG weight w: ε_guided = ε_uncond + w(ε_cond - ε_uncond); w=7-12 is standard for photorealistic image synthesis.
  • SDXL Turbo / LCM / Hyper-SD: distilled diffusion models using adversarial score distillation or consistency training achieve 1-4 step generation; these replace the iterative sampler with a learned shortcut from noise to image, trading sampler flexibility for speed.
  • Stable Diffusion 3.0 / FLUX: flow matching models parameterise the trajectory as straight paths x_t = (1-t) x_noise + t x_data; the ODE becomes dx/dt = x_data - x_noise (constant velocity), enabling highly efficient Euler integration with 20-28 steps; the sampler is trivially first-order with no need for DPM-Solver-class exponential integrators.
  • Video diffusion: temporal consistency requires joint space-time sampling; Stable Video Diffusion and Sora-class models use 3D U-Net or transformer architectures with the same DDIM/DPM-Solver samplers applied to space-time latents; motion prior conditioning replaces static text conditioning.
  • Inpainting and image editing via sampling: RePaint (Lugmayr et al. 2022) interleaves forward diffusion (adding noise to the unmasked region) with reverse sampling (denoising the full image), resampling the unmasked region at each step to enforce consistency; ControlNet (Zhang et al. 2023) adds spatial conditioning to the score network, enabling geometry-aware sampling.

LLM Decoding Strategy Comparison

StrategyVocabulary truncationAdaptive to model confidenceDistribution-preservingMain use case
GreedyArgmax (k=1)N/ANo (deterministic)Factual Q&A
Beam searchTop-k beamNoNo (MAP approx)Translation, summarisation
Top-kFixed kNoYes (within truncation)General generation
Top-p (nucleus)Dynamic p-nucleusPartiallyYesMost creative tasks
Min-pDynamic min-p floorYesYesCreative, instruction
TemperatureNone (redistributes)NoYes (at T=1)Combined with above
MirostatDynamic k via feedbackYesApproxLong-form generation
EtaEntropy-relative floorYesYesRobust generation
SpeculativeNone (exact)N/AYes (exact match)Throughput at scale
  • Beam search maintains K candidate sequences, extending each by all vocabulary tokens and pruning to the K highest cumulative log-probability beams; produces high-probability but often repetitive output; standard in neural machine translation (Sutskever et al. 2014) and summarisation (Lewis et al. 2020, BART).
  • Diverse beam search (Vijayakumar et al. 2018) adds a diversity penalty between beams to improve output variety; group beam search partitions beams into G groups encouraging diversity across groups.
  • Constrained beam search (Post & Vilar 2018) enforces the inclusion of specified lexical constraints (must-have terms) by maintaining a state machine over the beam; used in controlled generation and instruction following.

Stochastic Sampling in Reinforcement Learning from Human Feedback (RLHF)

  • RLHF fine-tuning (Ouyang et al. 2022, InstructGPT) uses the sampled outputs of the SFT model as the action space for PPO training: the LLM generates diverse responses via top-p sampling, a reward model scores each, and PPO updates the policy to maximise expected reward.
  • KL penalty: RLHF adds a KL divergence penalty KL(π_RL || π_SFT) to prevent the policy from collapsing to degenerate high-reward but low-quality outputs — this is itself a sampling constraint on the learned distribution.
  • DPO (Direct Preference Optimisation) (Rafailov et al. 2023) eliminates the reward model and PPO loop by directly fine-tuning on preference pairs, implicitly modelling RLHF as a change in the sampling distribution away from π_SFT proportional to exp(r/β).
  • Best-of-N sampling (also “rejection sampling fine-tuning”): generates N responses via temperature sampling, scores each with a reward model, returns the highest-scoring response; equivalent to IS sampling from π_SFT with reward weights but without policy update; effective for test-time compute scaling (Snell et al. 2024).
  • Speculative decoding meets RLHF: EAGLE-2 draft models must be trained on the RLHF-finetuned target model distribution to maintain the exact distribution guarantee; distribution shift between SFT and RLHF models requires retraining draft models.

Academic Context

  • The theoretical foundations of MCMC were established by physicists: Metropolis et al. (1953, Los Alamos) computed partition functions for interacting particle systems; Hastings (1970, Cambridge) generalised the acceptance rule.
  • The statistics community adopted MCMC with the Geman-Geman (1984) Gibbs sampler for image restoration and Gelfand-Smith (1990) for Bayesian computation, triggering the “MCMC revolution” enabling practical Bayesian inference.
  • Neal (1993, 2011) brought HMC into statistics; Hoffman & Gelman (2014) made it mainstream via NUTS and Stan; Girolami & Calderhead (2011) extended to Riemannian manifolds.
  • Robert & Casella’s “Monte Carlo Statistical Methods” (2004, 2nd ed.) remains the canonical graduate text, covering IS, MCMC, and variance reduction.
  • Chopin & Papaspiliopoulos’s “An Introduction to Sequential Monte Carlo” (2020, Springer) is the standard SMC reference.
  • Skilling’s nested sampling (2006) and its Cambridge implementations (PolyChord) produced widely used cosmological parameter estimation tools.
  • In LLM decoding: Holtzman et al. (2020) “The Curious Case of Neural Text Degeneration” (ICLR 2020, 3000+ citations) established nucleus sampling; Bradley et al. (2024) introduced min-p; Basu et al. (2021) introduced mirostat; Leviathan et al. (2023) and Chen et al. (2023) independently proposed speculative decoding.
  • In diffusion models: Song et al. (2020) derived DDIM; Lu et al. (2022) derived DPM-Solver; Karras et al. (2022) unified the design space — these three papers define the modern diffusion sampling landscape.
  • Convergence theory for MCMC: Rosenthal (1995) established quantitative convergence rates for finite-state Markov chains; Roberts & Tweedie (1996) extended geometric ergodicity to continuous state spaces; Hairer, Stuart & Vollmer (2014) proved spectral gaps for Langevin dynamics in high dimensions; these theoretical results underpin the practical convergence diagnostics used in Stan and PyMC.
  • Optimal transport and sampling: the Wasserstein distance W₂(p, q) measures the cost of transporting probability mass from p to q; Villani’s “Optimal Transport: Old and New” (2009) establishes that the optimal transport map T: p → q minimises quadratic transport cost; this geometric perspective motivates flow-based samplers, SVGD, and the Schrödinger bridge formulation of diffusion models.
  • Non-reversible MCMC: lifting tricks and Hamiltonian dynamics break the detailed balance condition while maintaining the stationary distribution, provably reducing autocorrelation time vs reversible chains; the Zig-Zag sampler (Bierkens et al. 2019) and bouncy particle sampler (BPS) are the leading non-reversible alternatives with proven spectral gap improvements.
  • Adaptive MCMC theory: Roberts & Rosenthal (2007) establish conditions (diminishing adaptation + containment) under which adaptive MH algorithms are ergodic despite history-dependent proposals — the theoretical foundation for Stan’s warm-up phase and PyMC’s step method adaptation.
  • Gelman-Rubin convergence and its limitations: the R-hat statistic is a between/within-chain variance ratio; Vats & Knudson (2021) prove that R-hat < 1.01 is necessary but not sufficient for convergence; their multivariate R-hat generalisation (using ESS for the joint chain vector) provides a stronger guarantee and is implemented in ArviZ v0.12+.
  • Information-geometric MCMC: the natural gradient Langevin algorithm (Patterson & Teh 2013) uses Fisher information metric preconditioning for SGLD, achieving significantly faster convergence than isotropic SGLD on deep neural network posterior inference; connections to mirror descent and Bregman divergences are explored in the natural MCMC literature.
  • Rare event sampling: metadynamics (Laio & Parrinello 2002) and adaptive biasing force methods add history-dependent bias potentials to MCMC to escape free-energy barriers — standard in computational chemistry and materials science for sampling rare conformational transitions; the bias is periodically removed via importance reweighting to recover unbiased ensemble averages.
  • Parallel tempering (replica exchange MC): runs M chains at different temperatures T₁ > T₂ > … > T_M and proposes swaps between adjacent temperature chains; the high-temperature chain rapidly explores the landscape while the low-temperature chain samples the target — widely used in protein folding and spin glass physics; implemented in parallel in MCX (GPU-accelerated photon transport).
  • Deep neural network posteriors: BNNs (Bayesian Neural Networks) with millions of parameters require novel MCMC approaches; Izmailov et al. (2021, SWAG — Stochastic Weight Averaging with Gaussians) approximates the BNN posterior with a low-rank Gaussian around the SGD trajectory; Wenzel et al. (2020) demonstrate that HMC produces better-calibrated posteriors than mean-field VI for moderately-sized ResNets on CIFAR-10.

Current Landscape (2026)

  • MCMC software: Stan (Columbia, ~200K active users, NUTS with adaptation); PyMC v5 (Python, 50K+ stars); NumPyro (JAX, hardware-accelerated NUTS); BlackJAX (DeepMind/INRIA, modular MCMC kernels composable in JAX); flowMC (Glasgow/UCL, normalising flows + MCMC for gravitational wave parameter estimation).
  • LLM inference frameworks (llama.cpp, vLLM, TGI, Ollama) expose temperature, top-p, top-k, min-p, and mirostat as first-class parameters.
  • Speculative decoding with EAGLE-2 draft models is now the default serving strategy for Llama-3.1-70B and Mistral-class models, delivering 2.5-3.5× throughput gains.
  • Diffusion model samplers are fully modularised in Node-Based Diffusion Pipeline Interface (30M+ users) and InvokeAI, with DPM-Solver++(2M) and Euler-a as defaults.
  • Research in 2025-2026 focuses on consistency models (Song et al. 2023) and flow matching (Lipman et al. 2022; SD 3.0) framing generation as straight-line ODE trajectories amenable to 1-4 step inference — potentially superseding multi-step diffusion samplers.
  • Approximate Bayesian Computation (ABC/likelihood-free inference) is growing for Simulation-Based Inference in systems biology and particle physics where likelihoods are intractable but forward simulators exist.
  • Diffusion models for discrete sequences: discrete DDPM (Austin et al. 2021, D3PM), absorbing diffusion, and masked diffusion models bring sampling-based generation to proteins, molecules, and code, requiring novel discrete ODE/SDE analogues.
  • Test-time compute scaling via sampling: OpenAI o1/o3 and DeepSeek R1 demonstrate that sampling multiple candidate solutions at test time and selecting the best (best-of-N) or aggregating via majority voting (self-consistency, Wang et al. 2023) substantially improves reasoning accuracy; the sampling temperature, N, and selection criterion are critical hyperparameters.
  • Structured prediction and constrained sampling: grammar-constrained decoding (Willard et al. 2023, Outlines library; llama.cpp GBNF grammars) enforces valid JSON/Python/SQL output by masking logits to zero for tokens violating the grammar at each step; equivalent to sampling from a conditional distribution p(xₜ | x_{<t}, grammar-valid) defined by a deterministic finite automaton.
  • Adaptive computation and early exit: speculative decoding is one instance of a broader adaptive sampling family where computation is allocated proportional to difficulty; LLM cascades route easy queries to small fast models and hard queries to large accurate models, using sampling-based uncertainty estimates (token entropy, ensemble disagreement) as routing signals.
  • Protein structure generation via diffusion sampling: AlphaFold3 and RoseTTAFold Diffusion (Baker Lab, UW) apply diffusion sampling to all-atom protein coordinate spaces; DDIM/DDPM samplers over SE(3) or torsion angle spaces generate diverse backbone conformations; RFdiffusion (Watson et al. 2023, Nature) produced novel protein binders validated experimentally in weeks vs years of rational design.
  • Drug discovery via molecular diffusion: DiffSBDD, TargetDiff, and DiffDock (Corso et al. 2023) apply SE(3)-equivariant diffusion samplers to molecular docking, generating diverse binding pose distributions rather than a single predicted pose — enabling ensemble docking for drug-lead prioritisation.
  • Sampling for AI safety: uncertainty quantification via MC dropout, deep ensembles, and MCMC posteriors is used to detect distribution shift and out-of-distribution inputs in safety-critical deployments (Lakshminarayanan et al. 2017); calibrated probability estimates from well-sampled posteriors are required for regulatory compliance under the EU AI Act’s high-risk AI system provisions.

UK Context (Imperial / Edinburgh / UCL / Cambridge / Manchester)

  • Cambridge Statistical Laboratory / Cavendish: John Skilling (Cavendish) developed nested sampling; Will Handley (Cavendish) leads PolyChord and applies it to Planck CMB data, ACT and SPT analyses.
  • MRC Biostatistics Unit Cambridge (BSU) (Dir: Sylvia Richardson): develops MCMC methods for genomic and clinical trial data, including trans-dimensional samplers for model uncertainty; their software informs NICE health technology assessment.
  • Cambridge Machine Learning Group (Zoubin Ghahramani, José Miguel Hernández-Lobato): sparse GP inference with MCMC, doubly stochastic variational inference, and flow-based normalising posterior approximations.
  • Edinburgh School of Informatics: Bayesian and Probabilistic ML group (Dirk Husmeier, Michael Gutmann) working on SMC and ABC for differential equation models in systems biology; Gutmann’s likelihood-free inference underpins simulation-based inference tools across the physical sciences.
  • Edinburgh’s Institute of Astronomy applies nested sampling (MultiNest, PolyChord) to exoplanet characterisation and stellar population inference.
  • Imperial College London: Girolami & Calderhead (2011) JRSS-B paper on Riemannian manifold HMC is one of the most cited MCMC papers of the 2010s; Imperial’s Centre for Process Systems Engineering applies LHS and QMC for uncertainty propagation in pharmaceutical manufacturing.
  • UCL Gatsby Computational Neuroscience Unit (Arthur Gretton, Mark van der Wilk): kernel-based hypothesis testing related to sample quality assessment (MMD, Stein discrepancy); Gaussian process-based surrogate-assisted sampling for expensive likelihoods.
  • UCL Statistics (Ricardo Silva, Ioanna Manolopoulou): SMC for causal inference and spatiotemporal models.
  • Alan Turing Institute (multi-site, London): funds Probabilistic Programming theme co-developing BlackJAX; national competence in uncertainty quantification for AI safety.
  • Manchester Mathematics (Kody Law): MCMC for PDE-constrained inverse problems and ensemble Kalman sampling for climate model calibration.
  • Northern England industrial: Hartree Centre (Daresbury, Cheshire, UKRI/STFC) applies QMC and LHS for industrial CFD and digital twin calibration; Rolls-Royce (Derby) and BAE Systems (Warton, Lancashire) use nested sampling and MCMC for engineering reliability and sensor fusion in autonomous systems.
  • Leeds and Sheffield: the Leeds Institute for Data Analytics applies MCMC for spatial health data modelling and infectious disease surveillance; Sheffield’s Department of Automatic Control and Systems Engineering (ACSE) uses particle filter (SMC) methods for industrial process monitoring and fault detection in steel and chemical manufacturing.
  • Newcastle and North East: Newcastle University’s Population Health Sciences Institute applies Bayesian hierarchical models with NUTS for multi-site clinical trial analysis; the NIHR Newcastle Biomedical Research Centre uses MCMC-based Bayesian adaptive trial designs for dose-finding studies in oncology.
  • Cambridge’s broader sampling ecosystem: beyond the MRC BSU and Cavendish, the Wellcome Sanger Institute (Hinxton, Cambridge) applies MCMC-based demographic inference (PSMC, MSMC) for ancient DNA population genetics; the European Bioinformatics Institute (EMBL-EBI) uses SMC-based phylogenetic inference (tsinfer, Kelleher et al. 2019) for whole-genome genealogy reconstruction from thousands of sequenced human genomes.
  • Imperial College Statistics: the Department of Mathematics, Statistics Section (Oliver Ratmann, Seth Flaxman) applies MCMC and ABC for HIV transmission network inference and COVID-19 epidemiological modelling; Imperial’s Bayesian phylogenetics group uses BEAST2 (MCMC-based Bayesian evolutionary analysis) for pathogen genomic surveillance, deployed during the Zika, Ebola, and COVID-19 outbreaks in collaboration with Public Health England / UKHSA.
  • UK National Computing Infrastructure for MCMC: ARCHER2 (UK national HPC, Edinburgh, 750K CPU cores) and Baskerville (Birmingham GPU cluster) support large-scale parallel MCMC for cosmological parameter estimation, genomics, and climate model calibration; the UKRI Probabilistic AI programme (2024-2029, £50M) funds sampling algorithm research across these institutions.
  • Edinburgh’s Machine Learning for Health Clinic applies Gaussian process-based active learning sampling strategies to NHS EHR data, adaptively selecting clinical records for manual review to improve diagnostic model training efficiency — combining active learning with MCMC posterior uncertainty quantification.
  • Manchester’s Henry Royce Institute for Materials Science applies nested sampling (NS-Materials, Partay et al.) and MCMC to crystallographic structure prediction from X-ray diffraction data — identifying low-energy crystal polymorphs relevant to pharmaceutical formulation.

Future Directions (2026-2030)

  • Consistency distillation and flow matching will likely replace multi-step diffusion samplers for image/video generation within 2-3 years, reducing NFE from 20-50 to 1-4 whilst matching quality.
  • Discrete diffusion models (D3PM, masked diffusion) extend denoising to categorical and structured data (proteins, molecules, code), requiring entirely new sampling strategies adapted to discrete probability simplices.
  • In LLM decoding: speculative decoding with fine-grained tree drafts and mixture-of-draft models will likely achieve 4-8× throughput gains; parallel decoding (Jacobi, lookahead) approaches breaking the autoregressive bottleneck are under active development; controlled generation via classifier guidance and RLHF-shaped sampling will be increasingly prominent as alignment requirements tighten.
  • For MCMC: gradient-informed transport maps (Cabezas et al. 2024) learn normalising flows to precondition the posterior, reducing effective dimensionality for NUTS; stochastic gradient MCMC (SGLD, SGHMC) will scale to billion-parameter posteriors for Bayesian deep learning; measure transport approaches (Marzouk MIT; Baptista ETH) reformulate sampling as deterministic pushforward problems solved by triangular maps.
  • Quantum sampling (quantum annealing via D-Wave, quantum Monte Carlo on NISQ hardware) remains speculative but is actively investigated for combinatorial optimisation applications where classical MCMC mixes poorly.
  • Neural posterior estimation (NPE, SBI toolkit) will increasingly replace likelihood-based MCMC for physical science inference, using amortised normalising flows trained on simulations.
  • Simulation-based inference (SBI) at scale: neural ratio estimation (NRE) and neural posterior estimation (NPE) using transformer-based summary statistic networks (BayesFlow, Radev et al. 2023) will enable Bayesian inference for models with 10M+ parameter simulators (climate, epidemiological, neuroscience ABM models).
  • Sampling for world models: video generation models (Sora, stable video) are increasingly framed as learned simulators; sampling from these models provides training data for downstream agents — MCMC-like correction of generated trajectories against real-world physics constraints is an open research direction.
  • Amortised MCMC: learning proposal distributions from data (neuralised MCMC, NeuTra-HMC, Hoffman et al. 2019; transport map MCMC) will reduce the number of gradient evaluations needed for convergence by orders of magnitude, enabling Bayesian inference at GPT-scale.
  • Diffusion samplers for combinatorial optimisation: score-based generative models applied to discrete graphs (DiffSolver, GFlowNets — Bengio et al. 2021) frame NP-hard combinatorial problems (TSP, protein folding, molecule generation) as sampling problems over exponentially large discrete spaces, with the diffusion sampler exploring the solution space.
  • Sampling under distributional constraints: constrained diffusion sampling (adding projected gradient steps to enforce conservation laws, physical constraints, or fairness criteria) will be essential for scientific generative models; energy-preserving Hamiltonian diffusion and gauge-equivariant lattice QCD samplers (Cranmer et al. 2023) demonstrate the principle.
  • Test-time scaling laws: empirical evidence from 2024-2025 (Snell et al. 2024; Brown et al. 2024) shows that allocating more sampling compute at inference (best-of-N, beam search with N=64-1024, self-consistency voting) produces accuracy improvements that follow power-law scaling in the sampling budget — parallel to training-time scaling laws; this drives architectural changes toward models optimised for sample efficiency rather than single-pass accuracy.

Research & Literature

  • [1] Metropolis N, Rosenbluth AW, Rosenbluth MN, Teller AH, Teller E (1953). Equation of state calculations by fast computing machines. J. Chem. Phys. 21(6):1087-1092.
  • [2] Hastings WK (1970). Monte Carlo sampling methods using Markov chains and their applications. Biometrika 57(1):97-109.
  • [3] Geman S, Geman D (1984). Stochastic relaxation, Gibbs distributions and the Bayesian restoration of images. IEEE TPAMI 6(6):721-741.
  • [4] Gelfand AE, Smith AFM (1990). Sampling-based approaches to calculating marginal densities. JASA 85(410):398-409.
  • [5] Duane S, Kennedy AD, Pendleton BJ, Roweth D (1987). Hybrid Monte Carlo. Physics Letters B 195(2):216-222.
  • [6] Neal RM (2011). MCMC using Hamiltonian dynamics. In: Brooks S et al. (eds) Handbook of Markov Chain Monte Carlo. Chapman & Hall/CRC.
  • [7] Hoffman MD, Gelman A (2014). The No-U-Turn Sampler: adaptively setting path lengths in Hamiltonian Monte Carlo. JMLR 15(1):1593-1623.
  • [8] Del Moral P, Doucet A, Jasra A (2006). Sequential Monte Carlo samplers. JRSS-B 68(3):411-436.
  • [9] Skilling J (2006). Nested sampling for general Bayesian computation. Bayesian Analysis 1(4):833-859.
  • [10] Robert CP, Casella G (2004). Monte Carlo Statistical Methods (2nd ed.). Springer, New York.
  • [11] Sobol IM (1967). On the distribution of points in a cube and the approximate evaluation of integrals. USSR Computational Mathematics and Mathematical Physics 7(4):86-112.
  • [12] McKay MD, Beckman RJ, Conover WJ (1979). Comparison of three methods for selecting values of input variables in the analysis of output from a computer code. Technometrics 21(2):239-245.
  • [13] Roberts GO, Gelman A, Gilks WR (1997). Weak convergence and optimal scaling of random walk Metropolis algorithms. Ann. Appl. Prob. 7(1):110-120.
  • [14] Neal RM (1998). Annealed importance sampling. Technical Report CRG-TR-98-1, Dept Computer Science, University of Toronto.
  • [15] Girolami M, Calderhead B (2011). Riemann manifold Langevin and Hamiltonian Monte Carlo methods. JRSS-B 73(2):123-214.
  • [16] Fan A, Lewis M, Dauphin Y (2018). Hierarchical neural story generation. ACL 2018. arXiv:1805.04833.
  • [17] Holtzman A, Buys J, Du L, Forbes M, Choi Y (2020). The curious case of neural text degeneration. ICLR 2020. arXiv:1904.09751.
  • [18] Basu S, Bhatt M, Parekh V (2021). Mirostat: a neural text decoding algorithm that directly controls perplexity. ICLR 2021. arXiv:2007.14966.
  • [19] Bradley E et al. (2024). Min-p sampling: diverse and coherent text generation. arXiv:2407.01082.
  • [20] Hewitt J, Manning CD, Liang P (2022). Truncation sampling as language model desmoothing. EMNLP Findings 2022. arXiv:2210.15191.
  • [21] Leviathan Y, Kalman M, Matias Y (2023). Fast inference from transformers via speculative decoding. ICML 2023. arXiv:2211.17192.
  • [22] Chen C, Borgeaud S et al. (2023). Accelerating large language model decoding with speculative sampling. arXiv:2302.01318.
  • [23] Song J, Meng C, Ermon S (2020). Denoising diffusion implicit models. ICLR 2021. arXiv:2010.02502.
  • [24] Lu C et al. (2022). DPM-Solver: a fast ODE solver for diffusion probabilistic model sampling. NeurIPS 2022. arXiv:2206.00927.
  • [25] Lu C et al. (2022). DPM-Solver++: fast solver for guided sampling of diffusion probabilistic models. arXiv:2211.01095.
  • [26] Karras T, Laine S, Aittala M, Hellsten J, Lehtinen J, Aila T (2022). Elucidating the design space of diffusion-based generative models. NeurIPS 2022. arXiv:2206.00364.
  • [27] Handley WJ, Hobson MP, Lasenby AN (2015). PolyChord: nested sampling for cosmology. MNRAS Letters 450(1):L61-L65.
  • [28] Chopin N, Papaspiliopoulos O (2020). An Introduction to Sequential Monte Carlo. Springer Series in Statistics.
  • [Supplementary A] Welling M, Teh YW (2011). Bayesian learning via stochastic gradient Langevin dynamics. ICML 2011. Foundational scalable MCMC reference.
  • [Supplementary B] Liu Q, Wang D (2016). Stein variational gradient descent. NeurIPS 2016. arXiv:1608.04471. Deterministic particle sampling via functional gradient flow.
  • [Supplementary C] Neal RM (2003). Slice sampling. Ann. Statist. 31(3):705-767. Systematic treatment of auxiliary-variable adaptive rejection.
  • [Supplementary D] Vehtari A, Gelman A, Simpson D, Carpenter B, Bürkner P-C (2021). Rank-normalization, folding, and localization: improved R-hat for assessing convergence. Bayesian Analysis 16(2):667-718.
  • [Supplementary E] Wang X, Wei J et al. (2023). Self-consistency improves chain of thought reasoning in language models. ICLR 2023. arXiv:2203.11171. Best-of-N / majority vote sampling for LLM reasoning.
  • [Supplementary F] Watson JL et al. (2023). De novo design of protein structure and function with RFdiffusion. Nature 620:1089-1100. Diffusion sampling for protein binder generation.
  • [Supplementary G] Bierkens J, Fearnhead P, Roberts G (2019). The Zig-Zag process and super-efficient sampling for Bayesian analysis of big data. Ann. Statist. 47(3):1288-1320. Non-reversible MCMC theory.
  • [Supplementary H] Cobb AD, Jalaian B (2021). Scaling Hamiltonian Monte Carlo inference for Bayesian neural networks with symmetric splitting. UAI 2021. BNN posterior sampling at scale.
  • [Supplementary I] Rezende D, Mohamed S (2015). Variational inference with normalizing flows. ICML 2015. arXiv:1505.05770. Flow-based posterior approximation bridging VI and MCMC.
  • [Supplementary J] Snell C et al. (2024). Scaling LLM test-time compute optimally. arXiv:2408.03314. Empirical scaling laws for sampling-based test-time compute.
  • [Supplementary K] Lipman Y et al. (2022). Flow matching for generative modelling. ICLR 2023. arXiv:2210.02747. Flow matching formulation enabling 1-step diffusion sampling via straight ODE trajectories.

Metadata

  • Domain correction: none. Page correctly classified as artificial-intelligence; the concept spans statistical computing, probabilistic modelling, LLM decoding, and diffusion model inference — all core AI subdisciplines.
  • Legacy term ID: AI-2047 (new assignment; no prior code in stub).
  • OWL axiom families: Compositional 8, Dependency 8, Capability 10, Implementation 12, Reduction 5, plus DisjointClasses (2), AnnotationAssertion (4), Property (5) = 54 total OWL statements.
  • Wikilink count: 71+ unique wikilinks across 11 relationship types.
  • Reference count: 28 numbered academic/industry references.
  • Enriched by: claude-sonnet-4-6, Phase 6 bulk run, 2026-05-17.

Provenance

  1. Robert CP, Casella G (2004). Monte Carlo Statistical Methods (2nd ed.). Springer.
  2. Holtzman A et al. (2020). The curious case of neural text degeneration. ICLR 2020. arXiv:1904.09751.
  3. Karras T et al. (2022). Elucidating the design space of diffusion-based generative models. NeurIPS 2022. arXiv:2206.00364.
  4. Bradley E et al. (2024). Min-p sampling: diverse and coherent text generation. arXiv:2407.01082.
  5. Hoffman MD, Gelman A (2014). The No-U-Turn Sampler. JMLR 15(1):1593-1623.
  6. Leviathan Y et al. (2023). Fast inference from transformers via speculative decoding. ICML 2023.
  7. Lu C et al. (2022). DPM-Solver. NeurIPS 2022. arXiv:2206.00927.
  8. Song J, Meng C, Ermon S (2020). DDIM. ICLR 2021. arXiv:2010.02502.
  9. Skilling J (2006). Nested sampling. Bayesian Analysis 1(4):833-859.
  10. Del Moral P, Doucet A, Jasra A (2006). Sequential Monte Carlo samplers. JRSS-B 68(3):411-436.
  11. Handley WJ et al. (2015). PolyChord. MNRAS Letters 450(1):L61-L65.
  12. Girolami M, Calderhead B (2011). Riemann manifold MCMC. JRSS-B 73(2):123-214.
  13. Chopin N, Papaspiliopoulos O (2020). Introduction to Sequential Monte Carlo. Springer.
  14. McKay MD et al. (1979). Latin Hypercube Sampling. Technometrics 21(2):239-245.
  15. Sobol IM (1967). Low-discrepancy sequences. USSR Comp. Math. 7(4):86-112.
  16. Metropolis N et al. (1953). J. Chem. Phys. 21(6):1087-1092.
  17. Hastings WK (1970). Biometrika 57(1):97-109.
  18. Geman S, Geman D (1984). IEEE TPAMI 6(6):721-741.
  19. Neal RM (1998). Annealed importance sampling. TR CRG-TR-98-1, U Toronto.
  20. Basu S et al. (2021). Mirostat. ICLR 2021. arXiv:2007.14966.
  21. Hewitt J, Manning CD, Liang P (2022). Eta/truncation sampling. EMNLP Findings 2022. arXiv:2210.15191.
  22. Chen C et al. (2023). Speculative sampling. arXiv:2302.01318.
  23. Lu C et al. (2022). DPM-Solver++. arXiv:2211.01095.
  24. Roberts GO, Gelman A, Gilks WR (1997). Optimal scaling RWM. Ann. Appl. Prob. 7(1):110-120.
  25. Fan A, Lewis M, Dauphin Y (2018). Hierarchical story generation (top-k). ACL 2018. arXiv:1805.04833.
  26. Neal RM (2011). MCMC using Hamiltonian dynamics. Handbook of MCMC. CRC Press.
  27. Xu Y et al. (2023). Restart sampling for improving generative processes. NeurIPS 2023. arXiv:2306.14878.
  28. MRC Biostatistics Unit Cambridge BSU — MCMC methods for health technology assessment. https://www.mrc-bsu.cam.ac.uk/.