Model optimisation is the set of techniques that reduce the size or computational cost of a trained model while preserving accuracy. It includes quantisation, pruning and distillation to make deployment more efficient.
Semantic Classification
Content
- Model optimisation transforms a trained model so that it runs faster, uses less memory or fits on constrained hardware. Quantisation lowers the numerical precision of weights and activations, pruning removes parameters with little effect, and distillation trains a smaller model to mimic a larger one.
- These techniques trade a small loss in accuracy for substantial gains in latency and cost, which matters for serving at scale and on edge devices. Optimised models feed inference engines that execute them on accelerators.