Adversarial training is a robustness technique that augments model training with adversarially perturbed examples generated to maximise the model’s loss. By solving an inner maximisation that crafts worst-case inputs within a bounded perturbation set and an outer minimisation over model parameters, it teaches models to resist adversarial attacks. It improves robustness against perturbations at the cost of additional computation and sometimes reduced clean accuracy.

Overview

  • Standard models can be fooled by tiny, carefully chosen perturbations that are imperceptible to humans; adversarial training directly defends against this failure mode.
  • It frames training as a saddle-point problem: an inner attacker finds the worst input within a bound, and the outer learner minimises loss against that attacker.
  • The technique is among the most reliable empirical defences, though it raises training cost and can trade off clean-data accuracy for robustness.

Mechanisms

  • Each batch is perturbed by an attack such as projected gradient descent, constrained to a small norm ball around the original input.
  • The model parameters are updated via Gradient Descent to correctly classify these adversarial examples.
  • This functions as a targeted form of Data Augmentation focused on the model’s current weaknesses.
  • Robustness is evaluated by attacking the trained model and measuring accuracy under attack, supporting Robustness guarantees.

Applications

  • Defending image and text Neural Network classifiers against evasion attacks.
  • Hardening safety-critical perception systems in autonomous platforms.
  • Improving reliability of models deployed in adversarial Security settings.
  • Studying generalisation and the relationship between robustness and Overfitting.

Provenance