Cost-efficient inference is the practice of serving model predictions at the lowest achievable compute and financial cost per request while meeting latency and quality targets. Techniques include dynamic batching, quantisation, caching, and context engineering to reduce token usage. It matters because inference, not training, dominates the lifetime cost of deployed large language models at scale.

Provenance