Triton Inference Server is NVIDIA’s open-source platform for serving machine-learning models in production across CPUs and GPUs. It supports multiple frameworks through a common interface, batches and schedules concurrent requests, and exposes models over HTTP and gRPC. Triton is a standard component of GPU-accelerated inference stacks, often paired with TensorRT-optimised models.
Overview
- Triton loads models from a versioned repository and serves them concurrently, applying dynamic batching to combine independent requests into efficient GPU kernels. It is framework-agnostic, hosting TensorRT, ONNX Runtime, PyTorch and TensorFlow backends behind one API, and supports model ensembles that chain pre-processing, inference and post-processing. Deployed on Kubernetes, it scales horizontally and exposes metrics for autoscaling and observability.
Key aspects
- Multi-framework model hosting behind a unified API
- Dynamic batching and concurrent model execution
- HTTP and gRPC inference endpoints
- Model ensembles and pre/post-processing pipelines
- Kubernetes-native scaling and metrics export
Applications
- GPU-accelerated production inference at scale
- Multi-model serving on shared accelerators
- Low-latency endpoints for recommendation and vision
- Autoscaled LLM and embedding inference services