Real-time inference is the execution of a trained machine-learning model to produce predictions within strict, low-latency time bounds suitable for interactive or streaming applications. It demands optimised serving infrastructure, efficient model formats, and often hardware acceleration to meet sub-second or millisecond response targets. Real-time inference enables responsive AI features such as recommendations, fraud scoring, and perception in autonomous systems.

Overview

  • Where batch inference tolerates minutes or hours, real-time inference must return results inside an interactive budget — typically milliseconds to a second.
  • Meeting that budget requires careful engineering: compact model formats such as ONNX, aggressive Model Optimization, and frequently hardware acceleration.
  • Serving systems keep models warm in memory, batch requests opportunistically, and route to accelerators to sustain throughput at low latency.
  • The same techniques extend to the edge, where On-Device Inference and Edge AI bring predictions close to the data source.

Key aspects

Mechanisms

Applications

  • Recommendation, search ranking and personalisation under interactive latency.
  • Fraud and risk scoring within transaction flows via Stream Processing.
  • Perception and control loops in robotics and autonomous systems.
  • Live analytics dashboards driven by Real-Time Analytics.

Provenance