Anomaly detection is a machine learning and statistical discipline concerned with identifying observations, sequences, or structural patterns that deviate significantly from a learned or assumed norm, signalling potential faults, threats, or novel phenomena. It operates across three principal modes: point anomaly detection (a single observation is outlying relative to the full dataset), contextual anomaly detection (an observation is anomalous given its local context, such as a transaction at an unusual time of day), and collective anomaly detection (a subsequence or group of observations is jointly anomalous relative to expected behaviour). The field draws on statistical modelling, machine learning, and signal processing to serve applications ranging from fraud detection and network intrusion detection to industrial fault monitoring, medical diagnostics, and log analysis.
Semantic Classification
Content
Compositional Relationships (Components)
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:hasPart ai:PointAnomalyDetection))
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:hasPart ai:ContextualAnomalyDetection))
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:hasPart ai:CollectiveAnomalyDetection))
Dependency Relationships
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:requires ai:StatisticalModelling))
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:requires ai:ThresholdCalibration))
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:requires ai:FeatureEngineering))
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:requires ai:DataPreprocessing))
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:dependsOn ai:PatternRecognition))
Capability Relationships
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:enables ai:FraudDetection))
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:enables ai:Cybersecurity))
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:enables ai:PredictiveMaintenance))
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:enables ai:IntrusionDetectionSystem))
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:enables ai:ModelMonitoring))
Implementation Relationships
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:implements ai:IsolationForest))
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:implements ai:OneClassSVM))
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:implements ai:LocalOutlierFactor))
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:implements ai:GaussianMixtureModel))
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:implements ai:Autoencoder))
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:implements ai:TransformerArchitecture))
Reduction Relationships
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:reducesTo ai:OutlierDetection))
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:reducesTo ai:NoveltyDetection))
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:reducesTo ai:FaultDetection))
Support Relationships
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:supports ai:NetworkSecurity))
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:supports ai:IoTSensorData))
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:supports ai:ModelMonitoring))
Usage Relationships
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:uses ai:GraphNeuralNetwork))
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:uses ai:VariationalAutoencoder))
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:uses ai:NormalisingFlows))
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:uses ai:LSTM))
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:uses ai:TransformerArchitecture))
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:uses ai:ActiveLearning))
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:uses ai:SHAPValues))
Parthood Relationships
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:partOf ai:MachineLearning))
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:partOf ai:StatisticalModelling))
Contrast Relationships
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:contrastsWith ai:SupervisedLearning))
SubClassOf(ai:AnomalyDetection
ObjectSomeValuesFrom(ai:contrastsWith ai:Classification))
About
Anomaly detection addresses the fundamental challenge of separating signal from noise in datasets where abnormal events are rare, ill-defined, or previously unseen. The discipline evolved from classical statistical work on outlier detection in the late nineteenth and early twentieth centuries — Grubbs’s test (1969) and the related Dixon’s Q test formalised single-outlier identification under Gaussian assumptions — through the machine learning revolution of the 1990s–2000s, which introduced density-based, proximity-based, and ensemble methods capable of handling high-dimensional, non-parametric distributions. The seminal survey by Chandola, Banerjee, and Kumar (2009, ACM Computing Surveys) systematised the field around the tripartite taxonomy of point, contextual, and collective anomalies that remains standard.
Three supervision paradigms define the practical design space. Supervised anomaly detection trains a classifier on labelled normal and anomalous examples; it achieves the highest precision when training labels are available and anomaly types are stable and well-enumerated, but fails entirely on novel anomaly types absent from the training distribution — a critical limitation in adversarial settings where attackers actively evolve their methods to evade detection. Semi-supervised anomaly detection trains only on normal data, learning a compact representation of normality and flagging departures; it handles novel anomaly types naturally and is the dominant paradigm in industrial and security settings where labelled anomalies are scarce and the definition of “anomaly” is “anything that doesn’t match the established normal.” Unsupervised anomaly detection requires no labels and infers normality directly from the data distribution; density estimation, clustering, and reconstruction-based methods all fall in this category. Hybrid approaches — including combinations of Autoencoder reconstruction scoring with Isolation Forest ensemble isolation scoring — achieve accuracy gains of 10–20 percentage points over single-method approaches in 2024–2025 benchmarks, with hybrid autoencoder-isolation forest systems reaching 0.98–0.99 accuracy on the CIC IOT-DIAD 2024 dataset.
The 2024–2026 period has been characterised by the emergence of Foundation Model-based anomaly detection as a new paradigm. Pre-trained Transformer Architecture models — particularly those trained on large time-series corpora — achieve state-of-the-art performance on multiple benchmark datasets without domain-specific Feature Engineering, establishing that general temporal representations transfer effectively to anomaly detection. AXIS (2025), a Large Language Model-based approach for Time Series anomaly detection, consistently ranks as a top performer on SED, TODS, UCR, and YAHOO benchmark datasets, outperforming domain-specific time-series foundation models on several evaluation metrics. TimeRadar (2026) demonstrates domain-rotation capability — adapting a single foundation model to anomaly detection across multiple industrial domains through lightweight fine-tuning — addressing the longstanding challenge of building generalisable anomaly detectors that do not require retraining from scratch for each new application domain.
Components / Architecture
Statistical and classical methods: Z-score and Grubbs test — parametric Point Anomaly Detection assuming Gaussian distributions; computationally trivial but sensitive to non-normality and multimodality. Gaussian Mixture Model (GMM) — density estimation over multi-modal distributions; anomalies are assigned low likelihood under the fitted mixture; EM optimisation converges to local optima and requires the number of components to be specified. DBSCAN — density-based clustering that identifies core, border, and noise points; points in low-density regions below a minimum-points threshold are flagged as outliers (noise); parameter selection (epsilon, MinPts) requires domain knowledge. Isolation Forest (Liu, Ting, and Zhou, 2008) — ensemble of random trees that isolate anomalies in fewer branch steps due to their sparsity relative to the bulk of the data distribution; average path length provides the anomaly score; efficient for large-scale, high-dimensional data (O(n log n) training complexity); scikit-learn implementations process approximately 50,000 samples per second on modern hardware. Local Outlier Factor (LOF) — proximity-based method comparing a point’s local reachability density to its neighbours’ densities; points with substantially lower density than their neighbours receive high LOF scores; effective for datasets with varying density regions but O(n²) naive implementation limits scalability. One-Class SVM — learns a minimum-volume hypersphere around normal data in kernel-induced feature space; rejects points outside the boundary with a margin controlled by the nu parameter; the RBF kernel with bandwidth selection via grid search is standard practice.
Deep learning methods: Autoencoder — an encoder-decoder neural network trained to reconstruct its normal training data with minimum reconstruction error; anomalies exhibit higher reconstruction error at inference time because the bottleneck representation fails to capture features outside the training distribution. Variational Autoencoder (VAE) — probabilistic extension that learns a latent distribution rather than a fixed latent code; anomaly scoring uses the evidence lower bound (ELBO), combining reconstruction quality with KL divergence from the prior; provides principled uncertainty estimates. Normalising Flows — learn exact density models via sequences of invertible transformations; anomalies receive low log-likelihood under the learned density; powerful but computationally intensive for high-dimensional data. LSTM and Temporal Convolutional Network (TCN) — model sequential dependencies in Time Series data; predict the next timestep and measure prediction error as the anomaly score; effective for seasonal and trending data where statistical baselines fail. Graph Neural Network (GNN) — represent multivariate sensor data or network traffic as graphs where edges encode relationships; detect structural anomalies in the graph or node-level deviations from learned normal relational patterns; particularly effective for Intrusion Detection System applications where lateral movement and unusual connection patterns in network traffic graphs constitute the anomaly signal. Transformer Architecture for anomaly detection — multi-head attention captures long-range temporal dependencies in Time Series data; masked reconstruction (training on reconstructing masked segments) forces the model to learn normal temporal patterns; Anomaly Transformer (Xu et al., 2022) introduced association discrepancy — the difference between attention weights derived from the raw series versus a learned prior — as an anomaly score, achieving state-of-the-art results on five time-series benchmark datasets.
Foundation models and LLM integration (2025–2026): STAR (2025) — State-aware Adapter for time series foundation models; uses lightweight adapter modules to specialise pre-trained time-series foundation models for anomaly detection without full fine-tuning; achieves 15–20% improvement over non-adapted foundation models on industrial benchmark datasets. AXIS (2025) — explainable anomaly detection using LLMs; provides natural-language explanations of flagged anomalies alongside anomaly scores, directly addressing the interpretability gap that limits operational deployment of black-box methods. TimeRadar (2026) — domain-rotatable foundation model for Time Series anomaly detection; a single pre-trained model adapts across manufacturing, network security, and finance domains through domain rotation adapters.
Operational pipeline components: Threshold Calibration — choosing the anomaly score cut-off to balance false-positive rate (alert fatigue — operators stop responding to alerts) against sensitivity (missed detections). Calibration is typically performed using a clean validation set of normal data to set the threshold at the 95th or 99th percentile of the normal score distribution. Business context drives the choice: a missed fraud transaction costing £10,000 justifies a very low threshold (high recall, more false positives reviewed by humans); an industrial monitoring alert that triggers a maintenance shutdown costing £50,000 per hour justifies a higher threshold. Active Learning — selectively querying human experts to label the highest-uncertainty or most informative anomaly candidates, iteratively improving the anomaly model while minimising labelling cost. Concept Drift adaptation — the normal distribution evolves over time (seasonal patterns change, legitimate user behaviour shifts, network topology changes); online adaptation and periodic retraining maintain model calibration; drift detectors (ADWIN, DDM, EDDM) monitor statistical properties of the anomaly score distribution and trigger retraining when drift is detected. SHAP Values and attention visualisation for interpretability — post-hoc explanation of anomaly scores by identifying which input features contributed most to the decision, enabling operators to understand and trust flagged alerts.
Use Cases / Major Families
Financial services — Fraud Detection: Payment fraud detection is the largest single industrial application of anomaly detection by transaction volume, processing hundreds of millions of transactions daily at global payment networks: Visa processes approximately 65,000 transactions per second at peak, with the anomaly detection system producing a risk score and optional hold or decline decision within a 100ms latency budget that must not impede the user experience of legitimate cardholders. UK authorised push payment (APP) fraud cost victims £450 million in 2024, reclassified as a national security risk and triggering mandatory reimbursement obligations under the PSR framework (effective October 2024) that directly increase banks’ financial exposure to fraud losses, creating strong incentives for investment in anomaly detection infrastructure. The operational constraint of sub-second inference latency constrains model complexity: gradient-boosted tree ensembles (XGBoost, LightGBM) dominate production real-time deployments due to their combination of high accuracy (AUPRC typically 0.85–0.95 on production payment data), fast inference (microsecond-scale tree traversal), and interpretability for regulatory reporting purposes. Graph Neural Network models operate asynchronously on transaction networks to detect money-laundering ring patterns — coordinated sequences across multiple accounts that are collectively anomalous but individually appear legitimate — with graph-level features encoding the topology of financial flows across accounts, beneficiaries, and intermediaries. Account takeover detection combines device fingerprinting (device ID, browser fingerprint, IP geolocation) with behavioural biometrics (typing rhythm, mouse movement, scroll patterns) and transaction pattern analysis to identify sessions where legitimate credentials are being used by a different person — a Contextual Anomaly Detection problem where the anomaly is not the transaction itself but the context in which it occurs. Market surveillance in equity and cryptocurrency markets deploys statistical process control and sequence anomaly detection to flag spoofing (placing and cancelling orders within milliseconds to manipulate price), wash trading (simultaneous purchase and sale between related parties to inflate volume and price), and unusual price movements indicative of insider trading or market manipulation; the FCA’s Market Oversight team uses ML-based surveillance in conjunction with human analysts reviewing flagged cases.
Cybersecurity — Intrusion Detection System and threat detection: Network anomaly detection identifies malicious traffic, lateral movement, and command-and-control (C2) communications by comparing real-time network behaviour against baselines of legitimate traffic established during a profiling period. Graph Neural Network models represent network hosts as nodes and communications as directed edges, detecting graph-structural anomalies — unusual connection patterns between hosts, novel communication paths to external IP ranges, unexpected protocol combinations — with performance substantially superior to packet-inspection classifiers on lateral movement scenarios: a 2025 Springer systematic review documented GNN-based detectors achieving AUC of 0.96–0.99 on UNSW-NB15 and CIC-IDS benchmark datasets versus 0.89–0.94 for random forest classifiers on the same benchmarks. User and Entity Behaviour Analytics (UEBA) builds statistical profiles of each user’s normal access patterns — which resources they access, at what times, from which devices and locations, at what data volumes — and flags sessions that deviate from the established profile; Isolation Forest and LSTM anomaly models are commonly deployed for UEBA given their ability to learn complex normal behaviour patterns without requiring labelled intrusion examples. Log anomaly detection applies sequence models (LSTM, Transformer Architecture) to system log streams (Windows Event Logs, Linux syslog, application logs, database audit logs) to surface unusual process chains, authentication failure sequences, and API call patterns; the HDFS and BGL log benchmark datasets are standard evaluation targets, and the DeepLog model (Du et al., SIGSAC 2017) established LSTM-based log anomaly detection as a practical alternative to rule-based SIEM correlation. The MITRE ATT&CK framework provides a structured taxonomy of adversary tactics (Initial Access, Execution, Persistence, Privilege Escalation, Defence Evasion, Credential Access, Discovery, Lateral Movement, Collection, Exfiltration, Command and Control, Impact) and techniques that guides what behaviour patterns anomaly detectors should prioritise surfacing, enabling alignment between anomaly detection engineering decisions and threat intelligence priorities. Adversarial robustness of anomaly detectors is a growing concern: attackers who understand the detection mechanism can craft adversarial examples that stay within the decision boundary by embedding malicious activity within the normal behaviour envelope or by gradually shifting the reference distribution through a series of small, individually innocuous changes that collectively move the baseline in an attacker-favourable direction.
Industrial and IoT — Predictive Maintenance and quality control: Bearing fault detection in rotating machinery (motors, turbines, compressors, pumps) from vibration signatures detects early-stage wear that manifests as changes in harmonic frequency content at integer multiples of the bearing pass frequency weeks or months before mechanical failure causes unplanned downtime. The vibration signal is typically sampled at 25–50 kHz, processed via Fast Fourier Transform to extract spectral features (bearing defect frequencies, sidebands, harmonics), and compared against reference spectra for the healthy machine; Autoencoder reconstruction error over the spectral features provides a continuous health index that drifts upward as faults develop. Production deployments demonstrate maintenance cost reductions of 10–20% and unplanned downtime reductions of 30–40% relative to scheduled time-based preventive maintenance, representing ROI periods of 12–24 months for typical industrial facilities at 2025 energy and downtime costs. Vision-based surface defect detection on production lines trains one-class models on images of defect-free products and flags reconstruction anomalies corresponding to scratches, cracks, voids, delaminations, and surface contamination without requiring an exhaustive defect catalogue; the MVTec AD dataset (Bergmann et al., CVPR 2019 — 5,354 images across 15 texture and object categories with pixel-level anomaly masks) is the primary benchmark, with 2025–2026 Transformer Architecture-based methods achieving mean per-class AUROC of 0.97–0.99 on MVTec AD. IoT Sensor Data monitoring applies streaming anomaly detection to high-frequency telemetry from manufacturing lines, smart grids, oil and gas pipelines, and building management systems; with internet-connected device populations projected to exceed 41 billion by 2026, the data generation rates in industrial IoT settings exceed the capacity for human review, making automated anomaly detection essential infrastructure rather than an optional enhancement. Edge deployment constraints — battery-powered IoT sensors, low-bandwidth communication, limited compute — motivate model compression and quantisation; scikit-learn Isolation Forest implementations processing 50,000 samples per second on modern edge hardware provide a practical baseline that satisfies real-time requirements while running on resource-constrained devices.
Healthcare: Clinical deterioration detection from vital signs streams (heart rate, blood pressure, SpO2, respiratory rate, temperature) in intensive care units deploys LSTM sequence models with anomaly scoring to provide early warning of sepsis, respiratory failure, cardiac arrest, and clinical deterioration hours before clinical staff would typically recognise the pattern from periodic manual observations; the National Early Warning Score 2 (NEWS2) system deployed across UK NHS hospitals provides a rules-based baseline against which ML-based anomaly detectors are evaluated, with validated studies showing ML models providing 2–6 hour earlier deterioration warnings compared to NEWS2 triggers. The NHS AI Lab’s AIDE framework mandates that clinical AI systems including deterioration detectors demonstrate equitable performance across age, sex, ethnicity, and socioeconomic patient subgroups before deployment — a requirement driven by documented biases in early commercial deterioration scoring algorithms that performed less accurately on Black and Asian patient populations due to training dataset imbalances. Medical imaging anomaly detection trains one-class models on large datasets of healthy tissue images (chest X-rays, retinal scans, brain MRIs, cardiac echo images) and flags reconstruction anomalies corresponding to tumours, lesions, and rare pathologies as outliers from the normal manifold; the NHS Lung Cancer Screening Programme deploys AI-assisted nodule detection that screens approximately 150,000 high-risk individuals annually as of 2025, flagging suspicious lung nodules for radiologist review — a system where the anomaly detector’s recall performance (sensitivity for true cancer nodules) is the safety-critical metric while precision determines the downstream radiologist workload. Genomics anomaly detection flags unusual mutation patterns, structural variants, copy-number variations, and rare germline variants deviating from population-level baselines in clinical sequencing pipelines; one-class learning is essential because the universe of pathogenic variants is both incompletely catalogued and continuously expanding as new variant-disease associations are discovered.
AI system monitoring: Model Monitoring for deployed machine learning systems uses anomaly detection across three distinct signal streams: input feature distributions (detecting when the distribution of features at serving time diverges from the training data distribution — Concept Drift and distributional shift that degrades model accuracy over time), output prediction distributions (detecting unusual patterns in model outputs that may indicate model malfunction, adversarial inputs, or systematic errors in a new data segment), and system performance metrics (detecting infrastructure anomalies including latency spikes, throughput degradation, memory leaks, and dependent service outages in ML serving pipelines). Population Stability Index (PSI) and Jensen-Shannon divergence over feature distributions provide scalar drift metrics that trigger retraining or rollback alerts when they exceed defined thresholds; ADWIN and Kolmogorov-Smirnov tests provide statistical hypothesis testing frameworks for distribution comparison. LLM output auditing applies anomaly scoring over generated text embeddings to flag outputs that deviate from the expected response distribution — a mechanism for detecting hallucinations (responses that confidently assert facts unsupported by the model’s context), unexpected topic shifts (the model’s output distribution shifts toward topics absent from its normal operational domain), and policy-violating content (outputs in the tail of the aligned response distribution that have not been filtered by safety classifiers). The intersection of anomaly detection and alignment monitoring is emerging as an important deployment practice: treating misaligned model outputs as anomalies relative to the expected aligned output distribution provides a complementary detection mechanism to explicit safety classifiers, potentially catching novel jailbreak patterns that safety classifiers have not been trained to recognise.
Satellite and space systems: Spacecraft telemetry anomaly detection monitors hundreds to thousands of housekeeping parameters (temperatures, voltages, currents, gyroscope readings, command counts) to detect equipment degradation, software faults, and environmental effects before they cause mission-critical failures. LSTM-based one-class models trained on nominal operational data detect deviations from expected telemetry patterns; the SMAP (Soil Moisture Active Passive) satellite anomaly dataset and MSL (Mars Science Laboratory) dataset, released by NASA/JPL, are standard benchmarks with 250+ annotated anomaly windows across 55 spacecraft channels. The challenge of temporal gaps in telemetry data (communication blackouts during orbital passes) and evolving normal behaviour (gradual component degradation, seasonal temperature variations, mission phase changes) requires adaptive anomaly models that account for expected long-term trends. Spacecraft-specific Autoencoder architectures (A Comparison of Deep Learning Architectures for Spacecraft Anomaly Detection, arXiv:2403.12864) systematically compare reconstruction-based methods across satellite telemetry data, demonstrating that variational autoencoders with temporal attention mechanisms achieve the best anomaly detection performance while providing calibrated uncertainty estimates for operator confidence scoring.
Environmental and climate monitoring: Environmental sensor networks monitoring air quality, water quality, soil conditions, and biodiversity indicators deploy anomaly detection to flag equipment malfunctions (sensor drift, clogging, calibration failure), unusual environmental events (pollution spikes, illegal dumping, extreme weather precursors), and long-term trend deviations (land surface temperature anomalies, precipitation deficits, species population collapses). The distinction between contextual anomalies (readings anomalous for the current season, weather conditions, or time of day) and genuine environmental events (readings that genuinely represent unusual environmental conditions rather than sensor malfunctions) requires domain-informed anomaly models that incorporate meteorological covariates. Newcastle University’s Urban Observatory, operating approximately 300 sensors monitoring air quality, traffic, energy consumption, and environmental conditions across Newcastle-upon-Tyne, provides one of Europe’s largest urban environmental sensor networks and has produced research on streaming anomaly detection for city-scale sensor data with real-time alert generation for environmental compliance monitoring.
Academic Context
Anomaly detection as a formal discipline emerged from statistics. Edgeworth’s (1887) treatment of outlier accommodation in regression established the philosophical precedent that statistical models should be robust to a small fraction of deviating observations that do not conform to the assumed data-generating process. Grubbs’s (1969) formalisation of the single outlier detection test under Gaussian assumptions provided the first widely-adopted hypothesis-testing framework for point anomaly detection; the test statistic is the maximum absolute deviation from the mean standardised by the sample standard deviation, and rejection of the null hypothesis (no outliers) at a specified significance level identifies the most extreme observation as an outlier. Dixon’s Q test (1951) provided an alternative for small samples. The transition to machine learning approaches was catalysed by the KDD Cup 1999 network intrusion detection competition, which provided 4.9 million labelled connection records across four attack categories (DoS, R2L, U2R, probe) and spurred a decade of competitive classifier development that established classification-based anomaly detection as the dominant industrial paradigm for network security — despite the conceptual limitation that classification requires exhaustive enumeration of attack types, which is unrealistic in adversarial settings where attackers continuously develop novel techniques.
Schölkopf et al.’s one-class SVM (2001, Neural Computation) and Breunig et al.’s Local Outlier Factor (2000, SIGMOD) established the principal classical non-parametric approaches that remain widely deployed in production systems. The one-class SVM uses the kernel trick to map normal training data to a high-dimensional feature space and finds the minimum-volume hypersphere enclosing a specified fraction (1-ν) of the training data; test points outside the hypersphere are flagged as anomalies. The computational complexity of kernel SVM scales as O(n²) to O(n³) in training, limiting its scalability to large datasets. LOF addresses the problem of datasets with heterogeneous density regions — where a global threshold would miss anomalies in low-density normal regions — by computing each point’s anomaly score relative to its k-nearest neighbours’ densities rather than against a global distribution. Liu, Ting, and Zhou’s Isolation Forest (2008, IEEE ICDM) provided the key conceptual innovation that anomalies are “few and different” — sparsely distributed in feature space — so they can be isolated in fewer splits of a random decision tree than normal points, which are densely distributed. The isolation score (average path length to leaf node across an ensemble of trees) provides an anomaly score that does not require density estimation and scales as O(t·n) in training where t is the number of trees (typically 100). The isolation principle has been extended to handle streaming data (iForestASD, 2018), high-dimensional data (Extended Isolation Forest, Hariri et al., 2021), and categorical features (Categorical Isolation Forest), establishing a family of practical tools for diverse industrial settings.
The Deep Learning era transformed anomaly detection through two successive waves. The first wave (2014–2019) adapted discriminative deep learning architectures to one-class learning: the autoencoder reconstruction paradigm (Hinton and Salakhutdinov, Science 2006 — foundational deep autoencoder work; Chalapathy et al., 2018 — deep one-class classification survey) trained encoder-decoder networks on normal data and used reconstruction error as an anomaly score. Ruff et al.’s Deep SVDD (2018, ICML) adapted the minimum enclosing hypersphere principle to deep representations by training a neural network to map normal training data into a compact region around a fixed centre in representation space, providing a deep alternative to kernel one-class SVM without the kernel trick’s computational cost. Kingma and Welling’s Variational Autoencoder (2013, ICLR 2014) provided the first principled probabilistic framework for deep generative anomaly scoring through the Evidence Lower BOund (ELBO), combining reconstruction fidelity and KL divergence from the prior into a unified anomaly score; subsequent work (SVAE, VQVAE-based anomaly detection) extended the VAE paradigm to more complex data distributions including images, time series, and graphs.
The second wave (2020–2026) was driven by attention mechanisms and Transformer Architecture models. The key insight of Xu et al.’s Anomaly Transformer (2022, ICLR) is that anomalies have distinctive attention patterns: normal time series points have concentrated self-attention to adjacent temporal neighbours (reflecting autocorrelation), while anomaly points have diffused attention to the whole series (because they do not exhibit the expected temporal correlation structure). The “association discrepancy” between Gaussian prior-guided attention weights and series self-attention weights provides a principled anomaly score that outperformed all prior methods on five benchmark datasets at time of publication. Transformer-based approaches have achieved state-of-the-art performance on MVTec AD for industrial visual inspection through masked image modelling pre-training followed by reconstruction-based anomaly scoring; the Anomaly Detection in the Wild (ADWIN) approach uses Vision Transformers pre-trained on ImageNet-21k as feature extractors with no domain-specific fine-tuning, demonstrating strong zero-shot transfer to industrial anomaly detection.
The current frontier (2025–2026) is defined by Foundation Model approaches. PatchTST (Nie et al., 2023, ICLR) demonstrated that pre-trained Transformer models with channel-independent patch-based tokenisation — treating each channel’s historical values as a sequence of 16-point patches similar to image patches in ViT — achieve strong zero-shot time-series anomaly detection without domain-specific training. TimesFM (Google, 2024), trained on a corpus of approximately 100 billion time points from public and proprietary sources, demonstrated that time-series foundation models achieve performance competitive with specialised methods on anomaly detection benchmarks. Sundial (2025) extended this to a family of highly capable time series foundation models, with the key innovation of using time-domain and frequency-domain representations jointly during pre-training to improve generalisation across anomaly detection applications with diverse spectral characteristics. AXIS (2025, arXiv:2509.24378) demonstrated that LLM-based anomaly detection with natural language explanation generation outperforms domain-specific methods on SED, TODS, UCR, and YAHOO benchmark datasets while providing interpretable reasons for flagged anomalies, addressing the critical deployability gap of black-box deep learning anomaly detectors in regulated and high-stakes settings. The ICLR 2025 track “Towards a General Time Series Anomaly Detector with Adaptive Bottlenecks and Dual Adversarial Decoders” and “Can LLMs Understand Time Series Anomalies?” highlight active debates about the limits of LLM pattern-matching versus genuine semantic understanding of anomaly context.
Key benchmark datasets define the evaluation landscape: KDD Cup 1999 (4.9 million network connection records; widely criticised for excessive replication and non-representative attack distribution but still used for reproducibility comparison), UNSW-NB15 (modern network traffic with 9 attack categories across 2.5 million records, designed to address KDD99 shortcomings), MVTec AD (5,354 images across 15 industrial categories with pixel-level masks; the dominant benchmark for visual anomaly detection with 400+ papers published against it since 2019), NAB (Numenta Anomaly Benchmark — 58 time-series streams from server metrics, IoT data, and financial data with anomaly windows rather than point labels to reflect real operational uncertainty), and YAHOO S5 (5,500 real and synthetic anomalous time series from Yahoo production traffic, widely used for evaluating streaming anomaly detectors). A persistent challenge is benchmark contamination: as methods are tuned against these datasets, apparent performance improvements may reflect dataset-specific optimisation rather than genuine generalisation. The call for living benchmarks that continuously add new real-world anomaly cases — analogous to the Common Vulnerabilities and Exposures (CVE) database for security vulnerabilities — is growing in the research community.
Current Landscape (2026)
The anomaly detection market is growing at a compound annual growth rate (CAGR) of 12.48% through 2025–2035, driven by expansion of cloud monitoring, IoT deployments, and regulatory requirements for fraud and intrusion detection. The global anomaly detection market is projected to reach USD 8.07 billion in 2026, growing to approximately USD 28.00 billion by 2034 (Precedence Research, 2025). The BFSI (banking, financial services, and insurance), government, healthcare, and IT services sectors are the largest adopters.
The most significant technical development of 2025–2026 is the integration of Foundation Model pre-training into anomaly detection pipelines. AXIS (2025) demonstrated that LLM-based approaches for Time Series anomaly detection with natural language explanations outperform domain-specific methods on multiple benchmark datasets while providing operator-readable explanations of flagged events — directly addressing the interpretability gap that limits operational deployment of black-box neural methods. The STAR framework (2025, arXiv:2510.16014) introduced State-aware Adapters that specialise pre-trained time-series foundation models for anomaly detection through lightweight fine-tuning, achieving 15–20% improvement over non-adapted foundation models on industrial benchmark datasets while preserving the zero-shot generalisation capability that makes foundation models valuable across diverse domains. The TimeRadar foundation model (2026, arXiv:2602.19068) achieves domain-rotatable generalisation through explicit domain-rotation adapter modules, enabling a single pre-trained model to detect anomalies across manufacturing, network security, and finance domains without full retraining; the domain rotation mechanism learns to shift the model’s internal representations to align with the statistical characteristics of each target domain, bridging the gap between general pre-training and domain-specific performance requirements. Incremental self-supervised learning approaches (ScienceDirect, 2025) apply Transformer Architecture models with online update mechanisms that continuously refine the normal behaviour model as new data arrives, addressing the Concept Drift problem without requiring periodic batch retraining by treating the updated model as the new reference for subsequent anomaly scoring. These foundation model approaches collectively represent a paradigm shift from the classical pipeline of domain-specific feature engineering → single-purpose model training → domain-specific deployment toward a unified pre-training → lightweight adaptation → cross-domain deployment pattern that dramatically reduces the engineering effort required for each new anomaly detection application.
Hybrid methods combining classical and deep learning approaches have become the production standard. A 2024 study on the CIC IOT-DIAD dataset demonstrated that Autoencoder and Isolation Forest hybrid models achieve 0.98–0.99 accuracy on IoT device anomaly detection, 18% better than either single-method approach. Industrial deployments of hybrid methods report 76% reduction in false positive rates compared to traditional threshold-based systems, substantially reducing alert fatigue in security operations centres and industrial monitoring control rooms. Scikit-learn implementations of Isolation Forest and related classical methods process approximately 50,000 samples per second on standard edge hardware, enabling real-time anomaly scoring in resource-constrained IoT deployments.
Concept Drift management has become a recognised operational discipline for deployed anomaly detection systems. As normal behaviour evolves — customer spending patterns change, network topology is reconfigured, production processes are modified — static anomaly models become miscalibrated, leading to either increasing false positive rates (if the normal distribution shifts toward the flagging boundary) or increasing false negative rates (if new fraud patterns are within the old normal region). Online learning frameworks that continuously update normal behaviour models, drift detectors (ADWIN, DDM) that trigger retraining when statistical properties of the anomaly score distribution change significantly, and ensemble methods that maintain multiple models trained at different historical time windows are standard approaches.
In the regulatory landscape, anomaly detection capabilities are referenced in multiple compliance frameworks. The EU AI Act (August 2026 high-risk requirements) requires post-market monitoring — effectively mandating deployment of Model Monitoring anomaly detection for high-risk AI systems. NIST SP 800-94 governs network-based anomaly detection for US government systems. IEC 62443 industrial cybersecurity standards require anomaly detection in industrial control system zone monitoring. PCI-DSS 4.0 (effective 2025) strengthens requirements for real-time transaction monitoring in payment systems. UK FCA’s Machine Learning in Financial Services guidance (updated 2024) requires model risk management including monitoring for distributional shift in AI-driven financial models.
UK Context
The UK anomaly detection landscape is shaped by three major industrial applications: financial fraud detection in the City of London, cybersecurity monitoring in government and critical national infrastructure, and NHS clinical deterioration detection. The financial fraud problem has reached national security scale: Authorised Push Payment (APP) fraud cost victims £450 million in 2024, prompting mandatory reimbursement obligations under the PSR (Payment Systems Regulator) framework effective October 2024, which in turn drives investment in real-time anomaly detection infrastructure across UK banks and payment processors. The UK Cyber Security Sectoral Analysis 2025 (DSIT) documents the UK cyber security industry at approximately £11.9 billion revenue, with anomaly detection and threat intelligence comprising a significant and growing fraction.
UK NHS anomaly detection deployments span clinical deterioration warning in acute trusts (National Early Warning Score 2 — NEWS2 — deployed nationally, with ML-augmented versions under evaluation), medical imaging anomaly detection in screening programmes (the NHS Lung Cancer Screening Programme’s AI-assisted nodule detection is one of the largest deployed radiological anomaly detection systems in the world, screening approximately 150,000 high-risk individuals annually as of 2025), and data quality monitoring in NHS Digital pipelines. High-profile NHS cybersecurity incidents (the June 2024 ransomware attack on Synnovis, a pathology services provider, caused approximately 3,000 outpatient appointments and 1,500 elective operations to be postponed) have driven investment in network anomaly detection and user behaviour analytics across NHS trusts.
Academic centres of anomaly detection research in the UK include: the University of Edinburgh’s Bayesian Machine Learning group, which has produced foundational contributions to Gaussian process-based change point detection (Adams and MacKay, 2007 — online change point detection using product partition models; Turner et al., 2009 — state space methods for non-parametric anomaly detection) that underpin probabilistic streaming anomaly detection methods widely deployed in financial and scientific monitoring applications; Imperial College London’s Data Science Institute, with research on adversarial robustness of anomaly detection systems (demonstrating that deep learning anomaly detectors trained on network traffic are vulnerable to adversarial perturbations that cause anomalous traffic to be classified as normal) and on causal inference methods for distinguishing genuine anomalies from confounded distributional shifts; the University of Manchester’s AI for Healthcare group, which investigates clinical deterioration detection from electronic health record signals, biosignal anomaly detection from wearable devices in remote patient monitoring, and EHR-based early warning for sepsis with a focus on equitable performance across NHS patient populations; Leeds University’s Data Analytics and Computational Statistics group, contributing theoretical foundations in functional data analysis for continuous-time anomaly detection relevant to industrial sensor streams; Newcastle University’s Urban Observatory — one of Europe’s largest urban sensor networks with approximately 300 sensors monitoring air quality, traffic flows, pedestrian counts, energy consumption, noise levels, and microclimatic conditions across Newcastle-upon-Tyne — which has published methods for managing sensor anomalies (distinguishing genuine environmental events from equipment malfunctions) and streaming anomaly detection for urban monitoring at the scale and heterogeneity of a real city; and Sheffield’s AMRC (Advanced Manufacturing Research Centre), which deploys vision-based anomaly detection for aerospace component quality inspection (surface defect detection on turbine blades, fuselage panels, and composite structures) and audio anomaly detection for machine health monitoring on multi-axis CNC machining centres, with results validated against destructive testing and non-destructive evaluation (NDE) ground truth.
The UKRI EPSRC project “FAIR” (Fault-Aware Integrated Robotics, University of Sheffield and BAE Systems, 2024–2027) develops anomaly detection for robotic manufacturing cells, combining vibration signal analysis with vision-based quality inspection. BT Group’s security operations team publishes annual threat intelligence reports drawing on anomaly detection over the UK’s largest commercial IP network, providing empirical data on UK-specific threat actor behaviour and detection performance.
Future Directions (2026–2030)
Foundation model consolidation and edge deployment: The trend toward Foundation Model approaches will consolidate the fragmented landscape of domain-specific anomaly detectors. Pre-trained Transformer Architecture models with adaptation layers will replace bespoke classical models in most standard application domains by 2028, following the pattern established in NLP and computer vision where foundation models displaced hand-engineered feature pipelines. The engineering challenge is deployment at the edge — resource-constrained IoT devices and industrial controllers with limited compute — requiring model compression, quantisation, and knowledge distillation from large foundation models to small inference-efficient models.
Explainable anomaly detection as a regulatory requirement: The EU AI Act’s requirement that high-risk AI system outputs be explainable, combined with sector-specific transparency obligations in financial services (FCA), healthcare (MHRA), and cybersecurity (NIS2 directive), will make explainability a mandatory component of anomaly detection systems rather than an optional enhancement. LLM-based explanation generation (as demonstrated in AXIS) and SHAP Values attribution over neural anomaly scores will become standard components of production anomaly detection pipelines. Standardised explanation formats analogous to structured anomaly reports (anomaly type, affected features, confidence interval, recommended human action) will emerge from regulatory guidance and industry standards bodies.
Continual learning and drift-resistant systems: As the pace of Concept Drift increases in adversarial settings (fraud attackers adapting to detection systems), static anomaly models will be replaced by continual learning systems that update their normality models online without catastrophic forgetting of historical normal patterns. Meta-learning approaches that rapidly adapt to distribution shifts from small samples, and out-of-distribution generalisation methods that maintain calibration across domain shifts, are active research priorities. The challenge of “catastrophic forgetting” in online-updated anomaly models — where rapid adaptation to new normal patterns causes forgetting of older patterns, potentially creating backdoors for historical attack patterns — motivates replay-based continual learning and elastic weight consolidation approaches.
Multimodal anomaly detection: Industrial deployments increasingly combine multiple sensor modalities (vibration + temperature + acoustic + visual) into unified anomaly models. Multimodal Foundation Model approaches trained on diverse sensor types will enable richer characterisation of normal operational conditions and more precise anomaly localisation, identifying not just that an anomaly exists but which sensor modality shows the strongest deviation and what physical phenomenon that deviation corresponds to. Multi-sensor fusion also provides robustness against single-sensor failure or adversarial manipulation of individual measurement channels.
Privacy-preserving federated anomaly detection: Distributed deployments across organisations that cannot share raw data — multiple banks sharing fraud pattern knowledge without sharing customer transaction records, multiple hospitals sharing clinical deterioration model updates without sharing patient records, multiple energy utilities sharing grid anomaly patterns without exposing operational security data — will adopt federated learning frameworks for collaborative anomaly model training. The UK Finance Fraud Insights initiative (UK Finance, 2024–2026) has pioneered federated anomaly detection across 20 member banks, enabling shared fraud pattern knowledge without centralising customer transaction data; initial results suggest 15–25% improvement in fraud detection AUPRC compared to each bank’s standalone models, driven by the ability to detect cross-bank fraud rings that are invisible to any individual institution. Differential privacy mechanisms (adding calibrated Laplace or Gaussian noise to shared model updates to prevent reconstruction of individual training examples with provable privacy guarantees) and secure multi-party computation protocols for aggregating model gradients without any participant seeing others’ raw gradients will become standard components of privacy-preserving anomaly detection infrastructure. The UK ICO’s guidance on privacy-preserving machine learning (updated 2025) specifically endorses federated learning and differential privacy as technical measures that can satisfy GDPR Article 25 “privacy by design and by default” requirements for ML systems processing personal data, creating a regulatory incentive for adoption of these techniques in anomaly detection systems that process transaction, healthcare, or communications data.
Human-AI collaboration in anomaly investigation: As automated anomaly detection systems achieve high recall at the expense of precision (producing many false positives for human review), the challenge shifts from detection to triage and investigation. Human-AI collaborative investigation interfaces — where the anomaly detector surfaces evidence, explanations, and related historical cases alongside the anomaly score, and the human analyst confirms or dismisses flagged events — will replace simple alert queues with structured collaborative investigation workflows. SHAP Values attribution at the feature level, attention weight visualisation for transformer-based models, and natural language explanation generation (as demonstrated in AXIS) will become standard components of anomaly detection user interfaces. The analyst’s investigation outcome — dismiss, escalate, confirmed anomaly — provides valuable labels for continuous model improvement through Active Learning, creating a virtuous cycle where analyst attention is directed to the highest-uncertainty cases and analyst decisions improve model calibration. Feedback integration through online learning updates, without catastrophic forgetting of historical normal patterns, is a key engineering challenge for this human-AI collaborative model that will be addressed through elastic weight consolidation, replay buffers, and progressive network architectures.
Research & Literature
- Chandola, V., Banerjee, A., and Kumar, V. (2009). Anomaly Detection: A Survey. ACM Computing Surveys, 41(3), 1–58.
- Goldstein, M. and Uchida, S. (2016). A Comparative Evaluation of Unsupervised Anomaly Detection Algorithms for Multivariate Data. PLOS ONE, 11(4).
- Liu, F.T., Ting, K.M., and Zhou, Z.-H. (2008). Isolation Forest. IEEE ICDM 2008, 413–422.
- Breunig, M.M., Kriegel, H.-P., Ng, R.T., and Sander, J. (2000). LOF: Identifying Density-Based Local Outliers. ACM SIGMOD 2000, 93–104.
- Schölkopf, B., Platt, J.C., Shawe-Taylor, J., Smola, A.J., and Williamson, R.C. (2001). Estimating the Support of a High-Dimensional Distribution. Neural Computation, 13(7), 1443–1471.
- Kingma, D.P. and Welling, M. (2014). Auto-Encoding Variational Bayes. ICLR 2014.
- Ruff, L., Vandermeulen, R., Goernitz, N., et al. (2018). Deep One-Class Classification. ICML 2018.
- Pang, G., Shen, C., Cao, L., and Hengel, A.V.D. (2021). Deep Learning for Anomaly Detection: A Review. ACM Computing Surveys, 54(2).
- Xu, J., Wu, H., Wang, J., and Long, M. (2022). Anomaly Transformer: Time Series Anomaly Detection with Association Discrepancy. ICLR 2022.
- Nie, Y., Nguyen, N.H., Sinthong, P., and Kalagnanam, J. (2023). A Time Series is Worth 64 Words: Long-Term Forecasting with Transformers. ICLR 2023.
- Chalapathy, R. and Chawla, S. (2019). Deep Learning for Anomaly Detection: A Survey. arXiv:1901.03407.
- Bergmann, P., Fauser, M., Sattlegger, D., and Steger, C. (2019). MVTec AD — A Comprehensive Real-World Dataset for Unsupervised Anomaly Detection. CVPR 2019.
- Lavin, A. and Ahmad, S. (2015). Evaluating Real-Time Anomaly Detection Algorithms — The Numenta Anomaly Benchmark. IEEE ICMLA 2015.
- Talagala, P.D., Hyndman, R.J., and Smith-Miles, K. (2021). Anomaly Detection in High-Throughput Data Streams. ACM TKDD.
- Morshedi, A., et al. (2025). A Comprehensive Review of Deep Learning Techniques for Anomaly Detection in IoT Networks. Wiley Engineering Reports. DOI:10.1002/eng2.70415.
- Kaur, A., et al. (2025). STAR: Boosting Time Series Foundation Models for Anomaly Detection through State-aware Adapter. arXiv:2510.16014.
- TimeRadar Collaboration (2026). TimeRadar: A Domain-Rotatable Foundation Model for Time Series Anomaly Detection. arXiv:2602.19068.
- AXIS Research Group (2025). AXIS: Explainable Time Series Anomaly Detection with Large Language Models. arXiv:2509.24378.
- Hybrid Anomaly Detection Study (2025). Hybrid Autoencoder and Isolation Forest for IoT Anomaly Detection. Engineering, Technology & Applied Science Research, 15(2).
- Alzantot, M., et al. (2025). The Role of GNNs, Transformers, and Reinforcement Learning in Network Threat Detection. Electronics (MDPI), 14(21):4163.
- Frontiers in AI (2025). A Deep Learning/Machine Learning Approach for Anomaly-Based Network Intrusion Detection. DOI:10.3389/frai.2025.1625891.
- NIST (2012). SP 800-94: Guide to Intrusion Detection and Prevention Systems. National Institute of Standards and Technology.
- IEC (2018). IEC 62443: Industrial Automation and Control Systems Security. International Electrotechnical Commission.
- Payment Systems Regulator (2024). APP Fraud Mandatory Reimbursement Policy. PSR Policy Statement PS24/1.
- DSIT (2025). UK Cyber Security Sectoral Analysis 2025. Department for Science, Innovation and Technology.
- Precedence Research (2025). Anomaly Detection Market Size to Hit USD 28.00 Billion by 2034. Market Research Report.
- BT Group (2025). Annual Threat Intelligence Report 2025. BT Security.
- UKRI EPSRC (2024). FAIR: Fault-Aware Integrated Robotics. Grant EP/X023869/1. Universities of Sheffield and BAE Systems.