A measure of disorder or uncertainty. In thermodynamics it quantifies the unavailable energy in a system, and in information theory it quantifies the average uncertainty or information content of a source.

Semantic Classification

Content

Compositional Relationships (Components)

SubClassOf(math:Entropy
  ObjectSomeValuesFrom(math:hasPart math:ShannonEntropy))
SubClassOf(math:Entropy
  ObjectSomeValuesFrom(math:hasPart math:DifferentialEntropy))
SubClassOf(math:Entropy
  ObjectSomeValuesFrom(math:hasPart math:JointEntropy))
SubClassOf(math:Entropy
  ObjectSomeValuesFrom(math:hasPart math:ConditionalEntropy))
SubClassOf(math:Entropy
  ObjectSomeValuesFrom(math:hasPart math:VonNeumannEntropy))
SubClassOf(math:Entropy
  ObjectSomeValuesFrom(math:hasPart math:RenyiEntropy))
SubClassOf(math:Entropy
  ObjectSomeValuesFrom(math:hasPart math:MinEntropy))
SubClassOf(math:Entropy
  ObjectSomeValuesFrom(math:hasPart math:RelativeEntropy))

Dependency Relationships

SubClassOf(math:Entropy
  ObjectSomeValuesFrom(math:requires math:ProbabilityTheory))
SubClassOf(math:Entropy
  ObjectSomeValuesFrom(math:requires math:RandomVariable))
SubClassOf(math:Entropy
  ObjectSomeValuesFrom(math:requires math:ProbabilityDistribution))
SubClassOf(math:Entropy
  ObjectSomeValuesFrom(math:depends_on math:Statistics))
SubClassOf(math:Entropy
  ObjectSomeValuesFrom(math:depends_on math:Logarithm))
SubClassOf(math:Entropy
  ObjectSomeValuesFrom(math:depends_on math:ExpectationOperator))

Capability Relationships

SubClassOf(math:Entropy
  ObjectSomeValuesFrom(math:enables math:DataCompression))
SubClassOf(math:Entropy
  ObjectSomeValuesFrom(math:enables math:RandomNumberGeneration))
SubClassOf(math:Entropy
  ObjectSomeValuesFrom(math:enables math:Cryptography))
SubClassOf(math:Entropy
  ObjectSomeValuesFrom(math:enables math:FeatureSelection))
SubClassOf(math:Entropy
  ObjectSomeValuesFrom(math:enables math:InformationGain))
SubClassOf(math:Entropy
  ObjectSomeValuesFrom(math:enables math:ChannelCapacity))
SubClassOf(math:Entropy
  ObjectSomeValuesFrom(math:supports math:MachineLearning))
SubClassOf(math:Entropy
  ObjectSomeValuesFrom(math:supports math:DeepLearning))
SubClassOf(math:Entropy
  ObjectSomeValuesFrom(math:supports math:ReinforcementLearning))
SubClassOf(math:Entropy
  ObjectSomeValuesFrom(math:supports math:BayesianInference))

Implementation Relationships

SubClassOf(math:Entropy
  ObjectSomeValuesFrom(math:implements math:CrossEntropyLoss))
SubClassOf(math:Entropy
  ObjectSomeValuesFrom(math:implements math:KLDivergence))
SubClassOf(math:Entropy
  ObjectSomeValuesFrom(math:implements math:MutualInformation))
SubClassOf(math:Entropy
  ObjectSomeValuesFrom(math:implements math:InformationBottleneck))
SubClassOf(math:Entropy
  ObjectSomeValuesFrom(math:implements math:EvidenceLowerBound))
SubClassOf(math:Entropy
  ObjectSomeValuesFrom(math:bridges_to math:Cryptography))
SubClassOf(math:Entropy
  ObjectSomeValuesFrom(math:bridges_to math:QuantumInformation))
SubClassOf(math:Entropy
  ObjectSomeValuesFrom(math:bridges_to math:Thermodynamics))

Reduction Relationships

SubClassOf(math:CrossEntropyLoss
  ObjectSomeValuesFrom(math:reducesTo math:Entropy))
SubClassOf(math:KLDivergence
  ObjectSomeValuesFrom(math:reducesTo math:Entropy))
SubClassOf(math:MutualInformation
  ObjectSomeValuesFrom(math:reducesTo math:Entropy))
SubClassOf(math:VonNeumannEntropy
  ObjectSomeValuesFrom(math:reducesTo math:ShannonEntropy))
SubClassOf(math:RenyiEntropy
  ObjectSomeValuesFrom(math:reducesTo math:ShannonEntropy))
SubClassOf(math:InformationGain
  ObjectSomeValuesFrom(math:reducesTo math:ConditionalEntropy))
SubClassOf(math:EvidenceLowerBound
  ObjectSomeValuesFrom(math:reducesTo math:KLDivergence))

About

Entropy is perhaps the only mathematical concept that appears, with the same formula, in the foundations of physics, the theory of communication, and the training of modern neural networks. Its history spans nearly 160 years and multiple scientific revolutions, from the steam engine to the internet to deep learning. Rudolf Clausius introduced entropy S in 1865 to formalise the observation that heat never spontaneously flows from cold to hot: dS = δQ_rev / T, where δQ_rev is the heat exchanged in a reversible process and T is the absolute temperature. The second law of thermodynamics states that the entropy of an isolated system never decreases; in practice, real processes are irreversible and total entropy always increases, establishing the “arrow of time” and setting fundamental limits on the efficiency of engines, refrigerators, and information-processing systems. Ludwig Boltzmann (1877) connected this macroscopic quantity to the microscopic world: S = k_B ln Ω, where Ω is the number of microstates consistent with the observed macrostate. This formula, inscribed on Boltzmann’s tombstone in Vienna, shows that entropy is a measure of multiplicity — a highly ordered, low-entropy state has few consistent microstates (a crystal), while a disordered, high-entropy state has many (a gas at equilibrium). The Boltzmann entropy reduces to the Gibbs entropy −k_B Σ p_i ln p_i when all states are not equally probable, providing a connection to statistical mechanics valid far beyond the simple counting argument.

Claude Shannon’s 1948 landmark paper “A Mathematical Theory of Communication” introduced the information-theoretic entropy H(X) = −Σ p(x) log₂ p(x), deriving it axiomatically as the unique function satisfying three reasonable properties: continuity in probabilities, maximum for the uniform distribution, and additivity for independent sources. Shannon immediately recognised that H was mathematically identical to Boltzmann–Gibbs entropy (up to a constant), and that this was not coincidental: both measure the same underlying notion of statistical uncertainty or unpredictability. Shannon’s source coding theorem established H(X) as the minimum average number of bits needed to represent symbols from source X without loss — a result of profound practical importance, establishing that compression algorithms have a hard theoretical ceiling. Shannon’s channel coding theorem established the channel capacity C = max I(X;Y) (maximum mutual information between input and output) as the theoretical maximum reliable communication rate, and proved that reliable communication is achievable at any rate below C. Together these two theorems created Information Theory as a field and provided the mathematical foundations for every digital communication and storage system in existence. The practical realisation of these limits took decades: Huffman coding (1952) approaches the source coding bound for single symbols; arithmetic coding (Rissanen, 1976) approaches it for sequences; turbo codes (Berrou et al., 1993) and LDPC codes (Gallager, 1960; MacKay and Neal, 1996) approach the channel capacity bound. Modern deep learning has found a new path: neural data compressors (Ballé et al., 2018) directly learn entropy-optimal representations by minimising the cross-entropy between a learned prior and the data distribution.

The connection between entropy and machine learning runs deep. Decision trees use information gain — the reduction in entropy achieved by a split — to select which feature to branch on at each node. The ID3 algorithm (Quinlan, 1986) built this approach into the canonical form of tree learning. The Naïve Bayes classifier is derived from maximum entropy principles. Logistic regression minimises cross-entropy loss. The entire field of probabilistic graphical models rests on the entropy-theoretic concept of Kullback-Leibler divergence D_KL(P||Q) = Σ P(x) log(P(x)/Q(x)), which measures the information cost of using distribution Q to approximate true distribution P. Variational inference — the foundation of modern approximate Bayesian methods including Variational Autoencoder networks — minimises a KL divergence between a tractable variational posterior q(z|x) and the true posterior p(z|x), maximising the Evidence Lower Bound (ELBO) = E_{q}[log p(x|z)] − D_KL(q(z|x)||p(z)). The KL term directly penalises the latent entropy of the encoder, regularising it toward the prior. Cross-entropy loss H(p,q) = −Σ p(x) log q(x) = H(p) + D_KL(p||q) is the standard training objective for virtually every modern classifier and sequence model: minimising cross-entropy loss with respect to model parameters q is equivalent to minimising KL divergence from the true data distribution p to the model distribution q, which is equivalent to maximum likelihood estimation. This deep equivalence underlies why cross-entropy loss works so well: it directly optimises the information-theoretic fit of the model to the data, and produces large, informative gradients even when the model’s predictions are far from correct.

Mathematical Formulations

The following formulas define the principal entropy variants in current use across Information Theory and Machine Learning:

  • Shannon Entropy (discrete): H(X) = −Σ_{x ∈ X} p(x) log_b p(x), where b=2 gives bits (the standard in CS), b=e gives nats (used in theoretical analysis and Variational Autoencoder literature), b=10 gives hartleys. Boundary convention: 0 log 0 = 0. Key properties: H(X) ≥ 0; H(X) = 0 iff X is deterministic (a single outcome has probability 1); H(X) ≤ log |X| with equality iff X is uniform over all |X| outcomes. Shannon proved these properties uniquely characterise H among continuous, symmetric, additive functions of probability distributions.

  • Differential Entropy (continuous): h(X) = −∫_{-∞}^{∞} f(x) log f(x) dx for a continuous random variable X with density f. Differential entropy can be negative (unlike discrete entropy) — the uniform distribution on [0, a] has h = log a, which is negative for a < 1. Differential entropy of the normal distribution N(μ, σ²) is h = (1/2) log(2πeσ²), used in Gaussian channel capacity calculations.

  • Joint Entropy: H(X,Y) = −Σ_{x,y} p(x,y) log p(x,y). Sub-additivity: H(X,Y) ≤ H(X) + H(Y), with equality iff X and Y are statistically independent.

  • Conditional Entropy: H(Y|X) = H(X,Y) − H(X) = −Σ_{x,y} p(x,y) log p(y|x). The chain rule: H(X₁, X₂, …, Xₙ) = Σᵢ H(Xᵢ | X₁, …, Xᵢ₋₁). Conditioning never increases entropy: H(Y|X) ≤ H(Y).

  • Mutual Information: I(X;Y) = H(X) + H(Y) − H(X,Y) = H(X) − H(X|Y) = D_KL(P(X,Y) || P(X)P(Y)). Symmetric, non-negative, zero iff X and Y are independent. Used in Feature Selection (mRMR, CMIM), information bottleneck, and as a channel capacity maximiser.

  • Cross-Entropy: H(p,q) = −Σ_x p(x) log q(x) = H(p) + D_KL(p||q). Standard neural network classification loss. When p is the one-hot true label distribution, H(p,q) = −log q(true_class), the negative log-likelihood of the correct class under the model.

  • Kullback-Leibler Divergence: D_KL(P||Q) = Σ_x P(x) log(P(x)/Q(x)). Non-negative (Gibbs inequality, equality iff P=Q); asymmetric (D_KL(P||Q) ≠ D_KL(Q||P) in general). The “information cost” of using Q when the true distribution is P. Central to Variational Autoencoder training (ELBO = reconstruction term − D_KL), to variational inference, and to model selection via minimum description length.

  • Von Neumann Entropy: S(ρ) = −Tr(ρ log ρ) for density matrix ρ. Zero for pure states (ρ = |ψ⟩⟨ψ|); maximised at log d for the maximally mixed state I/d in a d-dimensional Hilbert space. Reduces to Shannon entropy when ρ is diagonal. Central to Quantum Information theory: S(ρ_A) = S(ρ_B) for the reduced density matrices of a pure bipartite state quantifies entanglement.

  • Rényi Entropy: H_α(X) = (1/(1−α)) log Σ_x p(x)^α for order α > 0, α ≠ 1. Shannon entropy is the limit α → 1. Min-entropy H_∞(X) = −log max_x p(x) (the order-∞ limit) is used in cryptography as the operational entropy measure relevant to predicting the most likely outcome. Collision entropy H_2(X) = −log Σ p(x)² is used in randomness extractors and hash function analysis.

  • Topological and Permutation Entropy: Extensions used in dynamical systems analysis to characterise the complexity of time series, relevant to biomedical signal processing (EEG, ECG).

    Entropy in Machine Learning: Architecture and Training

    The information-theoretic perspective on machine learning originated with Good (1950), Jaynes (1957), and Rissanen (1978), but became dominant practice with the widespread adoption of neural network training by gradient descent. Every major component of modern deep learning involves entropy in a concrete computational role:

  • Cross-Entropy Loss for Classification: Given a K-class classifier with softmax output q = softmax(logits) and true one-hot label p, the cross-entropy loss L = −Σ_k p_k log q_k = −log q_{true_class}. Its gradient ∂L/∂logits = q − p (the difference between predicted and true distribution) is constant-magnitude regardless of the logit values, avoiding the gradient saturation that occurs with mean-squared error applied to sigmoid outputs. This is why cross-entropy is universally preferred for classification. For language model training (next-token prediction), cross-entropy is applied at every position, and the exponentiated per-token cross-entropy is the perplexity metric reported in NLP evaluations.

  • Information Gain in Decision Tree Learning: At each split node, information gain IG(feature f, dataset D) = H(D) − Σ_{v ∈ values(f)} (|D_v|/|D|) H(D_v). ID3 (Quinlan, 1986) selects the feature with maximum information gain; C4.5 normalises by the feature’s own entropy (gain ratio) to avoid bias toward features with many values. Random forests use either information gain or Gini impurity (a second-order approximation to entropy) for splitting. Feature Selection algorithms such as mRMR (Peng et al., 2005) and CMIM use mutual information to rank features by their information content about the target variable after removing redundancy.

  • Variational Autoencoders and the ELBO: The Variational Autoencoder (Kingma and Welling, 2014) trains an encoder q(z|x) (mapping inputs x to latent distributions z) and decoder p(x|z) (mapping latent codes to reconstructions) by maximising the Evidence Lower Bound: ELBO = E_{q(z|x)}[log p(x|z)] − D_KL(q(z|x)||p(z)). The first term rewards accurate reconstruction; the second term penalises the KL divergence between the encoder’s posterior and a standard normal prior, regularising the latent space. The entropy of the posterior H(q(z|x)) appears explicitly within D_KL: D_KL(q||p) = −H(q) + E_q[log(1/p)], so maximising ELBO is equivalent to maximising the differential entropy of the encoder distribution (within the KL constraint). This prevents posterior collapse to a point mass.

  • Information Bottleneck Theory: The information bottleneck (IB) method (Tishby, Pereira and Bialek, 2000) formalises representation learning as the problem of finding a compressed representation Z of input X that maximises mutual information I(Z;Y) about a target variable Y while minimising I(Z;X) (the amount of information retained from the input). The IB Lagrangian L = I(Z;Y) − β I(Z;X) trades off task relevance against compression, parameterised by β. Tishby and Schwartz-Ziv (2017) claimed that stochastic gradient descent causes deep networks to undergo two phases: initial fitting (increasing I(Z;Y)) followed by compression (decreasing I(Z;X)), suggesting that generalisation emerges through entropy compression. Saxe et al. (2019) showed this finding was highly dependent on the activation function and the mutual information estimator used; linear networks do not exhibit compression. The field remains active, with entropy-theoretic accounts of generalisation continuing to be developed alongside PAC-Bayes and flat minima perspectives.

  • Maximum Entropy Reinforcement Learning: Policy gradient methods (SAC, TRPO, PPO) add an entropy bonus H(π) to the expected return to prevent premature convergence to deterministic, suboptimal policies. The Soft Actor-Critic (SAC) algorithm (Haarnoja et al., 2018) optimises: J(π) = Σ_t E_{(s_t,a_t)~π}[r(s_t, a_t) + α H(π(·|s_t))], where α is a temperature parameter (automatically tuned in practice). Entropy-regularised policies are more robust, explore more thoroughly, and generalise better across tasks; SAC and its variants have become the dominant off-policy algorithms for continuous control in robotics as of 2025. The temperature α controls the trade-off between exploitation (maximising reward) and exploration (maximising entropy), analogous to the temperature parameter in thermodynamic Boltzmann distributions — another instance of the deep structural connection between physics and learning.

  • Maximum Entropy Models in NLP: Log-linear models trained by maximum entropy (MaxEnt) with feature constraints produce the least-biased distribution consistent with observed feature statistics. Conditional random fields (CRFs), used as decoding layers in Named Entity Recognition and sequence tagging, are maximum-entropy models over sequence labellings conditioned on input sequences. The maximum-entropy language model (Rosenfeld, 1996) was an important precursor to neural language models.

  • Feature Selection via Mutual Information: Mutual information I(feature; target) measures the information that a feature carries about the prediction target, capturing nonlinear dependencies that correlation cannot detect. mRMR (Minimum Redundancy Maximum Relevance, Peng et al., 2005) selects features that maximise I(feature; target) while minimising I(feature; other selected features), producing compact, informative feature subsets. Empirically, MI-based selection consistently outperforms Pearson correlation and F-statistic selection for nonlinear classification tasks.

  • Perplexity in Language Model Evaluation: Perplexity PP(LM, test_set) = 2^{H(test_set, LM)} = exp(average cross-entropy per token). A lower perplexity indicates that the language model assigns higher probability to the test data — it is less “surprised” by what it sees. Perplexity is the standard evaluation metric for language models, tracking the entire history from early n-gram models (PP ~ 100–200 on PTB) through LSTMs (PP ~ 60–70) to current large language models (GPT-4: PP < 10 on many benchmarks).

    Entropy in Cryptography and Security

    High entropy is a prerequisite for secure cryptographic key material. A key drawn uniformly from {0,1}^K has K bits of Shannon entropy; an adversary with no side information requires 2^{K-1} expected guesses to find it. If the key is derived from a low-entropy source — a poorly seeded pseudorandom number generator, a human-chosen password, or a hardware RNG with flawed physical noise collection — the effective security is reduced to the entropy of the seed, regardless of the key length. NIST SP 800-90B defines entropy estimation procedures for hardware random bit generators, specifying eight approved statistical tests that characterise the entropy per sample of a noise source. The Linux kernel’s /dev/random and /dev/urandom interfaces accumulate entropy from hardware interrupt timing, disk seek timing, and explicit hardware RNG inputs (RDRAND on Intel, RNDR on ARM) to provide a high-entropy source for cryptographic applications. The blocking behaviour of early /dev/random (blocking when estimated entropy drops below 64 bits) was a long-standing source of application-layer delays; the Linux 5.4 kernel changed /dev/random to be non-blocking, accepting that the ChaCha20 CSPRNG seeded at boot provides sufficient security.

    Shannon (1949) proved that the one-time pad (OTP) achieves perfect secrecy — the ciphertext is statistically independent of the plaintext — if and only if the key is drawn uniformly and independently of the message. This requires key entropy ≥ plaintext entropy: the entropy budget of the key must exceed that of the message. Shannon also introduced the concept of unicity distance U = H(K) / R (where H(K) is key entropy and R is the redundancy of the natural language), the minimum ciphertext length beyond which a brute-force search is expected to yield exactly one consistent decryption. This quantifies the practical security margin of historical ciphers. In Cryptography, min-entropy H_∞(X) = −log max_x p(x) is the operationally relevant entropy measure for key generation and randomness extraction: it captures the probability that an adversary correctly guesses the most likely value on a single attempt, and is used in the theory of randomness extractors and hash-based key derivation functions (HKDF, PBKDF2, scrypt, Argon2).

    Use Cases and Major Application Domains

  • Data Compression: Huffman coding (Huffman, 1952) constructs prefix-free codes that achieve expected code length within 1 bit of H(X) per symbol. Arithmetic coding (Rissanen and Langdon, 1979) approaches H(X) to within a fraction of a bit. Lempel-Ziv-Welch (LZW) and its variants (used in ZIP, gzip, PNG) achieve the entropy rate of stationary ergodic sources asymptotically. Modern learned image compression systems (Ballé et al., 2018; Minnen et al., 2018) use deep neural networks to learn entropy models optimised for real image data, outperforming JPEG 2000 and BPG at equivalent perceptual quality, with entropy coding remaining a core component.

  • Machine Learning Loss Functions: Cross-entropy loss is the universal training objective for probabilistic classifiers (logistic regression, softmax classifiers), language models (GPT, BERT, T5, Claude, Llama), segmentation models, and object detectors with classification heads. All major deep learning frameworks (PyTorch, JAX, TensorFlow, MXNet) implement cross-entropy as a first-class, numerically-stabilised operation via log-sum-exp tricks.

  • Natural Language Processing: Language model perplexity (exponentiated cross-entropy) tracks progress across the entire history of NLP, from n-gram models through LSTMs to transformer-based models. Entropy is also used in active learning for NLP — selecting the sentence or document with maximum predictive entropy for annotation most efficiently reduces model uncertainty per annotation effort. Text summarisation systems use entropy-based relevance scoring to identify information-dense sentences.

  • Cryptography and Security: Entropy pools, entropy estimation (NIST SP 800-90B), key derivation functions (HKDF, PBKDF2, Argon2), password strength metrics (NIST SP 800-63B uses entropy as the core model for password guessability), and physical unclonable functions (PUFs) all use Shannon or min-entropy as their foundational security metric.

  • Reinforcement Learning: Maximum entropy RL (SAC, MIRL, energy-based policies) is the dominant paradigm for continuous control tasks in robotics, simulation, and game-playing. The entropy bonus prevents mode collapse and encourages the discovery of diverse, robust strategies. SAC is the standard baseline for MuJoCo, DeepMind Control Suite, and RoboSuite benchmarks as of 2025.

  • Biomedical Signal Processing: Approximate entropy (ApEn, Pincus, 1991) and sample entropy (SampEn, Richman and Moorman, 2000) quantify the regularity and complexity of physiological time series (EEG, ECG, HRV, EMG, fMRI BOLD). Reduced entropy in cardiac signals indicates pathological regularity (arrhythmia, heart failure); reduced EEG entropy indicates loss of consciousness, anaesthetic depth, or epileptic seizure; increased entropy in gene expression profiles indicates tumour heterogeneity. These measures are used clinically and in research at UK NHS centres, including Leeds General Infirmary cardiology and Newcastle Hospitals ICUs.

  • Physics and Engineering: The second law of thermodynamics governs energy conversion: the efficiency of heat engines is bounded by the Carnot efficiency η_Carnot = 1 − T_cold/T_hot, derivable from entropy arguments. Information erasure has a thermodynamic cost: the Landauer limit k_B T ln 2 per erased bit sets the fundamental minimum energy for computation, connecting entropy to the physical limits of computing hardware. Non-equilibrium thermodynamics uses entropy production rate σ = dS/dt − Q_flow/T to characterise dissipation in biological systems, nanomachines, and active matter.

    Academic Context

    Entropy’s intellectual history spans multiple scientific communities and disciplinary revolutions. Clausius (1865) named the quantity from the Greek trope (transformation), capturing the directionality of thermodynamic processes. Boltzmann (1877) and Gibbs (1902) established its statistical mechanical foundation. The mathematician Norbert Wiener (1948), writing in “Cybernetics”, independently developed an information-theoretic entropy concept closely parallel to Shannon’s, emphasising its connection to feedback, control, and biological systems. Shannon’s 1948 Bell System Technical Journal paper remains one of the most cited papers in all of science, routinely credited as the founding document of the digital age. Edwin Jaynes (1957) formalised the maximum entropy principle as a general approach to Bayesian inference: the prior should be the maximum-entropy distribution consistent with known constraints, the least-informative assumption compatible with the evidence. Jaynes’s MaxEnt principle provided a unified framework connecting Thermodynamics, statistics, and information theory that remains influential in Bayesian deep learning and probabilistic ML. The mutual information perspective on feature selection and learning was developed by Battiti (1994), Peng, Long and Ding (2005 — mRMR), and Kwak and Choi (2002 — CMIM). The information bottleneck framework (Tishby et al., 2000) sparked a wave of research attempting to characterise deep learning generalisation in information-theoretic terms, generating ongoing debate about whether compression is a cause or byproduct of good generalisation. The theory of quantum entropy was built by von Neumann (1932), extended by Araki and Lieb (1970, strong subadditivity), and applied to quantum communication by Schumacher (1995) and Holevo (1998 — Holevo bound on classical information in quantum channels). The quantum entropy research community is particularly active at the University of Cambridge Cavendish Laboratory, University of Oxford Department of Computer Science, Imperial College London Physics, and UCL Physics.

    Primary journals: IEEE Transactions on Information Theory; Entropy (MDPI, open access, interdisciplinary); Journal of Statistical Physics; Physical Review Letters (thermodynamics and quantum entropy papers). Primary conferences: IEEE International Symposium on Information Theory (ISIT); Annual Allerton Conference on Communication, Control and Computing; NeurIPS (for entropy-based ML methods); ICML.

    Current Landscape (2026)

    In mid-2026, entropy remains one of the most intensively applied mathematical concepts across AI, computing, and physics. Cross-entropy loss is the near-universal training objective for large language models (GPT-4, Claude 3.5, Llama 3, Gemini Ultra), vision transformers, diffusion model discriminators, and multimodal embedding models. The information bottleneck perspective has been partially rehabilitated: Goldfeld and Polyanskiy (2020) proved that with Gaussian noise injected into the activations, compression does occur in the hidden layers; the debate has shifted to whether this compression is necessary for generalisation or an incidental byproduct of stochastic optimisation in overparameterised models. Maximum entropy Reinforcement Learning (SAC and its successors) became the dominant algorithm for real-world robot control in the 2024–2026 period, with implementations deployed in warehouse automation (Amazon Robotics), robotic surgery (Intuitive Surgical), and autonomous vehicle control (Waymo, Mobileye). In Cryptography, NIST’s finalised post-quantum standards (ML-KEM/Kyber, ML-DSA/Dilithium, SLH-DSA, 2024) required careful entropic analysis of lattice-based constructions, particularly their key generation procedures and NTT-based sampling algorithms. Von Neumann entropy and Rényi entropies are central to characterising quantum entanglement and quantum advantage on near-term quantum hardware, with quantum entropy estimators designed for NISQ devices appearing in arXiv preprints through 2025 (Quantum Neural Estimation of Entropies, arXiv:2307.01171). The MDPI Entropy journal’s special issue “Quantum Information and Probability: From Foundations to Engineering III” attracted significant 2025 submissions connecting quantum entropy to machine learning. The preprint “Information Physics of Intelligence” (arXiv:2511.19156, 2025) proposed a unified framework connecting Kolmogorov complexity, Shannon entropy, and Boltzmann thermodynamic entropy to characterise the fundamental limits on intelligent systems, generating wide discussion in both the physics and AI communities.

    UK Context

    The United Kingdom’s engagement with entropy and information theory is deep, historically grounded, and continues at the frontier of 2025–2026 research. Alan Turing’s wartime work at Bletchley Park used statistical methods functionally equivalent to information-theoretic analysis — “banburismus” was an iterative Bayesian procedure that computed log-likelihood ratios of Enigma key hypotheses, essentially measuring the discriminative entropy between competing models — and Turing corresponded directly with Shannon after the war, exchanging ideas about intelligence, computing, and information. I. J. Good, who worked with Turing at Bletchley, published “Probability and the Weighing of Evidence” (1950), one of the earliest UK texts connecting Bayesian inference and information-theoretic measures, and subsequently contributed to the theory of smoothing estimators (Good-Turing smoothing) that use entropy-based reasoning to estimate the probability of unseen events in language corpora.

    The University of Cambridge’s Statistical Laboratory (home of R. A. Fisher, who contributed to the maximum likelihood foundations of statistical entropy) remains a centre for information-theoretic statistics; Cambridge’s Cavendish Laboratory hosts quantum information research where von Neumann and entanglement entropy are key metrics, as part of the UK Quantum Technology Programme. David MacKay FRS (Cambridge), before his untimely death in 2016, was one of the world’s foremost information theorists; his open-access textbook “Information Theory, Inference, and Learning Algorithms” (2003, https://www.inference.org.uk/itprnn/book.pdf) is widely used in UK universities and by self-taught practitioners globally. Imperial College London’s Department of Computing has active groups in probabilistic ML and Bayesian deep learning where KL divergence and entropy regularisation are central; Imperial’s quantum computing group (part of the London Centre for Nanotechnology, in partnership with National Physical Laboratory in Teddington) investigates quantum entropy estimation and quantum error correction. The National Physical Laboratory (NPL) in Teddington contributes to UK entropy standards for random number generation, closely aligned with NIST SP 800-90B.

    The University of Edinburgh’s School of Informatics is a leading UK centre for probabilistic models in Machine Learning, particularly variational Bayes methods; the EPSRC Centre for Doctoral Training in Data Science (Edinburgh) trains researchers in information-theoretic methods including entropy in generative models. The University of Manchester’s Department of Computer Science applies entropy-based Feature Selection in industrial machine learning pipelines, with links to the Hartree Centre (STFC, Daresbury) for large-scale HPC-accelerated entropy computations. The University of Sheffield’s Department of Computer Science maintains the GATE NLP framework and applies entropy-based active learning to annotate large biomedical and legal corpora efficiently. In the NHS, sample entropy (SampEn) and approximate entropy (ApEn) are used in cardiology research at Leeds General Infirmary, the Freeman Hospital in Newcastle, and the Royal Brompton Hospital in London to characterise cardiac signal complexity, predict arrhythmia onset, and guide device therapy decisions in implanted defibrillator programming.

    Future Directions (2026-2030)

  • Quantum Entropy Estimation on Near-Term Hardware: Efficient quantum circuits for estimating von Neumann entropy and Rényi entropies on NISQ devices without full state tomography, enabling real-time characterisation of quantum system complexity and entanglement. The National Quantum Computing Centre (NQCC) at Harwell, operational from 2024, and IBM’s quantum network partnerships with UK universities provide access to hardware for such experiments. Quantum neural entropy estimators (arXiv:2307.01171) are an early step in this direction.

  • Information-Theoretic Foundations of LLM Generalisation and Scaling: Understanding why large language models generalise so remarkably well from Minimum Description Length and rate-distortion perspectives, and whether scaling laws can be derived from entropy-based arguments. This direction may provide principled guidance for model compression, distillation, and efficient deployment — connecting fundamental theory to industrial practice.

  • Thermodynamic Computing and the Landauer Limit: Physical computing implementations that operate at or near the Landauer limit (k_B T ln 2 per bit erasure), using entropy-aware circuit design to minimise energy dissipation. UK initiatives including EPSRC Quantum Engineering, ARM research at Cambridge, and Graphcore (Bristol-based AI chip company) are exploring reversible and near-reversible computing paradigms motivated by entropy economics.

  • Entropy-Aware Neural Architecture Search: Using information gain and mutual information as primary objective functions in neural architecture search (NAS) to find network structures that naturally compress input entropy in a task-optimal manner, rather than relying solely on accuracy or efficiency proxies.

  • Differential Privacy and Entropy Rate: The formal connection between differential privacy guarantees (ε, δ) and entropy rates of randomisation mechanisms is an active theoretical area. Future work will design privacy mechanisms that achieve tight entropy-privacy trade-offs, directly relevant to NHS data sharing, federated learning across NHS trusts, and the UK’s National Data Strategy.

  • Information Bottleneck for Foundation Models: Extending IB theory to vision-language foundation models, multimodal transformers, and retrieval-augmented generation architectures to understand how information about different modalities and retrieved documents is integrated, compressed, and routed through attention layers — with the goal of designing more efficient, robust, and interpretable architectures.

  • Entropy in AI Regulation and Auditing: The EU AI Act and emerging UK AI Liability Framework are beginning to require explainability and auditability for high-risk AI decisions. Entropy-based uncertainty quantification — measuring the predictive entropy H(y|x) of a model’s output as a calibrated confidence measure — will be a key component of AI auditing systems, providing regulators and auditors with principled measures of a model’s epistemic and aleatoric uncertainty.

    Research and Literature

    1. Clausius, R. (1865). “Ueber verschiedene für die Anwendung bequeme Formen der Hauptgleichungen der mechanischen Wärmetheorie.” Annalen der Physik, 125, 353–400. [Original thermodynamic entropy definition.]
    2. Boltzmann, L. (1877). “Über die Beziehung zwischen dem zweiten Hauptsatze der mechanischen Wärmetheorie und der Wahrscheinlichkeitsrechnung respektive den Sätzen über das Wärmegleichgewicht.” Sitzungsberichte der Akademie der Wissenschaften Wien, 76, 373–435. [S = k_B ln Ω.]
    3. Gibbs, J. W. (1902). Elementary Principles in Statistical Mechanics. Yale University Press. [Gibbs entropy H = −k_B Σ p_i ln p_i.]
    4. Shannon, C. E. (1948). “A mathematical theory of communication.” Bell System Technical Journal, 27(3), 379–423; 27(4), 623–656. [Founding paper of information theory; Shannon entropy; source and channel coding theorems.]
    5. von Neumann, J. (1932). Mathematische Grundlagen der Quantenmechanik. Springer. [Quantum entropy S(ρ) = −Tr(ρ log ρ).]
    6. Shannon, C. E. (1949). “Communication theory of secrecy systems.” Bell System Technical Journal, 28(4), 656–715. [Perfect secrecy; unicity distance; entropy-based cryptography.]
    7. Jaynes, E. T. (1957). “Information theory and statistical mechanics.” Physical Review, 106(4), 620–630. [Maximum entropy principle.]
    8. Kullback, S. & Leibler, R. A. (1951). “On information and sufficiency.” Annals of Mathematical Statistics, 22(1), 79–86. [KL divergence.]
    9. Huffman, D. A. (1952). “A method for the construction of minimum-redundancy codes.” Proceedings of the IRE, 40(9), 1098–1101. [Huffman coding approaching H(X).]
    10. Rissanen, J. (1978). “Modeling by shortest data description.” Automatica, 14(5), 465–471. [Minimum Description Length (MDL).]
    11. Schumacher, B. (1995). “Quantum coding.” Physical Review A, 51(4), 2738–2747. [Quantum source coding theorem; coined ‘qubit’.]
    12. Good, I. J. (1950). Probability and the Weighing of Evidence. Griffin. [UK foundational text; Bayesian entropy-based inference.]
    13. MacKay, D. J. C. (2003). Information Theory, Inference, and Learning Algorithms. Cambridge University Press. https://www.inference.org.uk/itprnn/book.pdf [Definitive UK-authored graduate text; freely available online.]
    14. Cover, T. M. & Thomas, J. A. (2006). Elements of Information Theory (2nd ed.). Wiley-Interscience. [Standard graduate textbook.]
    15. Quinlan, J. R. (1986). “Induction of decision trees.” Machine Learning, 1(1), 81–106. [Information gain in ID3.]
    16. Tishby, N., Pereira, F. C. & Bialek, W. (2000). “The information bottleneck method.” Proceedings of the 37th Allerton Conference on Communication, Control and Computing, 368–377. https://arxiv.org/abs/physics/0004057
    17. Tishby, N. & Schwartz-Ziv, R. (2017). “Opening the black box of deep neural networks via information.” arXiv:1703.00810. https://arxiv.org/abs/1703.00810
    18. Saxe, A. M. et al. (2019). “On the information bottleneck theory of deep learning: The linear case.” ICLR 2019.
    19. Haarnoja, T. et al. (2018). “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.” ICML 2018, 1861–1870. https://arxiv.org/abs/1801.01290
    20. Kingma, D. P. & Welling, M. (2014). “Auto-encoding variational Bayes.” ICLR 2014. https://arxiv.org/abs/1312.6114
    21. Peng, H., Long, F. & Ding, C. (2005). “Feature selection based on mutual information: Criteria of max-dependency, max-relevance, and min-redundancy.” IEEE Transactions on Pattern Analysis and Machine Intelligence, 27(8), 1226–1238. [mRMR.]
    22. Pincus, S. M. (1991). “Approximate entropy as a measure of system complexity.” Proceedings of the National Academy of Sciences, 88(6), 2297–2301. [ApEn for biomedical time series.]
    23. Richman, J. S. & Moorman, J. R. (2000). “Physiological time-series analysis using approximate entropy and sample entropy.” American Journal of Physiology, 278(6), H2039–H2049. [SampEn.]
    24. Quantum Neural Estimation Team (2023). “Quantum neural estimation of entropies.” arXiv:2307.01171. https://arxiv.org/pdf/2307.01171
    25. Information Physics of Intelligence Team (2025). “Information physics of intelligence: Unifying logical depth and entropy under thermodynamic constraints.” arXiv:2511.19156. https://arxiv.org/html/2511.19156v3
    26. MDPI Entropy Special Issue (2025). “Quantum Information and Probability: From Foundations to Engineering III.” https://www.mdpi.com/journal/entropy/special_issues/267RXFS74C
    27. NIST SP 800-90B (2018). “Recommendation for the entropy sources used for random bit generation.” National Institute of Standards and Technology. https://csrc.nist.gov/publications/detail/sp/800-90b/final
    28. arXiv 2407.12288 (2024). “Information-theoretic foundations for machine learning.” https://arxiv.org/pdf/2407.12288

    Key Terminology

  • Entropy H(X): Average uncertainty or information content of a random variable X; in bits (log base 2), nats (log base e), or hartleys (log base 10). Bounded: 0 ≤ H(X) ≤ log |X|.

  • Cross-Entropy H(p,q): Average number of bits needed to encode events from true distribution p using optimal code for approximating distribution q; standard neural network classification and language model training loss.

  • KL Divergence D_KL(P||Q): Information cost of using Q to approximate P; non-negative, asymmetric, zero iff P = Q. Central to variational inference and generative model training.

  • Mutual Information I(X;Y): Reduction in uncertainty about X when Y is known; equals H(X) − H(X|Y). Used in Feature Selection, information bottleneck, and active learning.

  • Information Gain: IG = H(parent) − Σ (weighted child entropy); entropy reduction achieved by a split in Decision Tree learning. Equivalent to mutual information between the feature and the target variable.

  • Perplexity: 2^H; exponentiated average cross-entropy per token; standard language model evaluation metric. Lower perplexity = better model.

  • Von Neumann Entropy: S(ρ) = −Tr(ρ log ρ); quantum analogue of Shannon entropy for density matrices; zero for pure states, maximum log d for maximally mixed d-dimensional states.

  • Differential Entropy h(X): Continuous-variable extension of Shannon entropy; can take negative values; used in Gaussian channel capacity calculations and in the VAE literature.

  • Information Bottleneck: The trade-off between I(Z;X) (compression of input into representation Z) and I(Z;Y) (predictive power of Z for target Y); a principled framework for representation learning.

  • Maximum Entropy Principle: Choose the probability distribution with maximum entropy subject to known constraints; due to Jaynes (1957); the least-biased assumption consistent with the available evidence.

  • Min-Entropy H_∞(X): −log max_x p(x); the operationally relevant entropy measure for cryptographic key generation and randomness extraction; lower-bounds all Rényi entropies.

  • Negentropy: J(X) = H(X_Gaussian) − H(X); measure of non-Gaussianity used in Independent Component Analysis (ICA) to find statistically independent signal sources.

  • Thermodynamic Entropy: S = k_B ln Ω (Boltzmann) or S = −k_B Σ p_i ln p_i (Gibbs); macroscopic measure of disorder or unavailable energy; governed by the second law of thermodynamics (dS/dt ≥ 0 for isolated systems).

  • Landauer Limit: k_B T ln 2; the minimum energy dissipated per bit of information erased in a computational operation; a fundamental thermodynamic constraint on computing hardware derived from entropy arguments.

Provenance