Inference Optimisation encompasses techniques and processes for reducing the computational cost, latency, and memory footprint of deploying trained machine learning models at runtime. Methods include quantisation, pruning, knowledge distillation, and hardware-specific kernel fusion. The goal is to make model inference faster and more efficient without significantly degrading predictive accuracy.