MLPerf is a suite of standardised benchmarks, governed by the MLCommons consortium, that measures the performance of machine learning hardware, software and systems for both training and inference. It defines fixed reference models, datasets, quality targets and submission rules so that results from different vendors are reproducible and directly comparable. MLPerf has become an industry reference for evaluating accelerators, frameworks and end-to-end systems on representative deep learning workloads.
- MLPerf is a suite of standardised Benchmark tests, governed by MLCommons, that measures Machine Learning system performance for both Training and Inference. It fixes reference models, datasets and quality targets so vendor results are reproducible and comparable.
Overview
- Before MLPerf, claims about accelerator and framework speed were hard to compare because vendors used different models, batch sizes and stopping criteria.
- MLPerf imposes a Benchmark Standard: defined tasks, target accuracies, and strict submission rules separating closed-division comparability from open-division innovation.
- Results cover datacentre and edge categories and report metrics such as time-to-train and queries-per-second, making the suite a touchstone for evaluating Hardware Accelerator and system designs.
Key aspects
- Reference models and datasets across vision, language and recommendation.
- Quality targets that submissions must reach before timing counts.
- Separate training and inference benchmark categories.
- Closed and open divisions balancing comparability and freedom.
- Audited submission process ensuring Reproducibility.
Applications
- Comparing GPUs and accelerators on representative Deep Learning workloads.
- Guiding procurement and capacity decisions for ML infrastructure.
- Validating framework and compiler optimisations on real models.
- Tracking industry progress in Throughput and Latency over time.