A model compression technique that reduces the numerical precision of neural network weights and activations from floating-point (FP32, FP16) to lower-bit integer representations (INT8, INT4, binary), decreasing memory footprint and improving inference speed on hardware with integer arithmetic units. Approaches include post-training quantisation and quantisation-aware training.

Semantic Classification

Content

Neural Network Quantisation (AI-0435) — content pending enrichment.

Provenance