Triton Inference Server is NVIDIA’s open-source platform for serving machine-learning models in production across CPUs and GPUs. It supports multiple frameworks through a common interface, batches and schedules concurrent requests, and exposes models over HTTP and gRPC. Triton is a standard component of GPU-accelerated inference stacks, often paired with TensorRT-optimised models.

Overview

  • Triton loads models from a versioned repository and serves them concurrently, applying dynamic batching to combine independent requests into efficient GPU kernels. It is framework-agnostic, hosting TensorRT, ONNX Runtime, PyTorch and TensorFlow backends behind one API, and supports model ensembles that chain pre-processing, inference and post-processing. Deployed on Kubernetes, it scales horizontally and exposes metrics for autoscaling and observability.

Key aspects

  • Multi-framework model hosting behind a unified API
  • Dynamic batching and concurrent model execution
  • HTTP and gRPC inference endpoints
  • Model ensembles and pre/post-processing pipelines
  • Kubernetes-native scaling and metrics export

Applications

  • GPU-accelerated production inference at scale
  • Multi-model serving on shared accelerators
  • Low-latency endpoints for recommendation and vision
  • Autoscaled LLM and embedding inference services

Provenance