Pruning is a model compression technique that removes redundant or low-importance parameters from a neural network to reduce its size and computational cost while preserving accuracy. It ranges from unstructured pruning of individual weights to structured pruning of whole neurons, channels or attention heads. Pruned models are typically fine-tuned to recover any lost accuracy and can be deployed with lower memory and latency.
- Pruning is a Model Compression technique that removes redundant or low-importance parameters from a Neural Network to reduce its size and computational cost while preserving accuracy.
- It spans unstructured pruning of individual Model Weights and structured pruning of whole neurons, channels or attention heads.
- Pruned models are usually fine-tuned to recover any lost accuracy.
- The result is lower memory and latency, aiding deployment to constrained targets such as Edge AI.
Overview
- Trained neural networks are typically over-parameterised, containing many weights that contribute little to predictions. Pruning exploits this redundancy to shrink the model.
- Unstructured pruning zeroes out individual weights according to importance criteria such as magnitude, yielding sparse matrices that need specialised kernels to accelerate.
- Structured pruning removes entire structural units such as channels, filters, neurons or attention heads, producing a dense smaller model that runs faster on standard hardware without custom sparse support.
- Pruning is commonly iterative: prune a fraction, fine-tune to recover accuracy, and repeat, often guided by schedules that gradually increase sparsity.
Key aspects
- Importance scoring of weights or structures.
- Unstructured versus structured granularity.
- Sparsity level and the accuracy-efficiency trade-off.
- Fine-tuning or retraining to restore performance.
- Hardware support required to realise speed-ups from sparsity.
Mechanisms
- Magnitude-based and gradient-based importance criteria.
- One-shot versus iterative prune-and-fine-tune schedules.
- The lottery-ticket hypothesis and sparse subnetwork discovery.
- Combination with quantisation and distillation for compound compression.
Applications
- Compressing large language models for cheaper inference.
- Deploying vision and speech models on edge and mobile devices.
- Reducing energy and memory footprint in production serving.
- Enabling TinyML on microcontrollers and embedded hardware.