Adversarial machine learning is the study of attacks that exploit vulnerabilities in machine learning models and the development of defences against them, encompassing threats across the model lifecycle including training-time data poisoning, evasion attacks at inference time, model inversion, and membership inference. Attackers craft carefully perturbed inputs or manipulate training data to cause misclassification, extract sensitive information, or degrade model performance, whilst defenders develop robust training procedures, certified defences, and detection mechanisms. The field spans both offensive security research and the development of trustworthy AI systems.
Content
- The field was crystallised by Szegedy et al.’s 2013 paper demonstrating that deep neural networks could be fooled by small, human-imperceptible perturbations to images—so-called adversarial examples. Goodfellow et al.’s Fast Gradient Sign Method (FGSM, 2015) provided an efficient algorithm for generating such perturbations and proposed adversarial training as a defence. Carlini and Wagner’s C&W attack (2017) subsequently broke many proposed defences, establishing the pattern of an arms race between attack and defence that characterises the field. Concurrently, Tramèr et al. demonstrated model stealing via black-box API queries, extending the threat model beyond white-box access.
- Technically, adversarial attacks are formalised as constrained optimisation problems: find a perturbation δ such that the perturbed input x+δ causes the model to output an incorrect label, subject to a constraint on the magnitude of δ under an Lp norm. Gradient-based attacks (PGD, AutoAttack) are most powerful in the white-box setting; transfer-based and query-efficient attacks address black-box deployments. Training-time attacks include clean-label poisoning, backdoor trojan injection (e.g., BadNets), and gradient-matching attacks that influence the model without modifying labels. Certified defences using randomised smoothing provide provable accuracy guarantees under L2-bounded perturbations.
- Applications of adversarial machine learning research extend across security-critical domains: adversarial patches that defeat object detection in autonomous vehicles and surveillance systems, adversarial perturbations that evade malware classifiers, poisoning attacks on federated learning systems, and face recognition evasion. Conversely, the field produces defensive infrastructure deployed in production AI systems including certified robustness layers, input preprocessing defences, and anomaly detectors that flag adversarially suspicious inputs.
- By 2024–2025, adversarial machine learning has expanded its scope to include large language models, multimodal systems, and agentic AI. Prompt injection attacks against LLMs—where adversarial instructions are embedded in user inputs or retrieved documents—have emerged as a critical threat to deployed AI assistants. Jailbreaking, indirect prompt injection, and multi-modal adversarial attacks (adversarial images that mislead vision-language models) are active research areas. Regulatory frameworks including the EU AI Act and NIST AI RMF explicitly address adversarial robustness requirements for high-risk AI systems, elevating the field from academic research to compliance obligation.