An AI model inference engine is the software runtime that executes a trained model to produce predictions from new inputs. It manages computation graph execution, hardware acceleration and memory to run models efficiently.
Semantic Classification
Content
Compositional Relationships (Components)
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:hasPart ml:KVCache))
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:hasPart ml:ContinuousBatching))
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:hasPart ml:SpeculativeDecoding))
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:hasPart ml:OperatorFusion))
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:hasPart ml:PrefixCaching))
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:hasPart ml:FlashAttention))
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:hasPart ml:RequestScheduler))
Dependency Relationships
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:requires ml:GPU))
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:requires ml:HardwareAccelerator))
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:requires ml:ModelWeights))
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:requires ml:ModelOptimization))
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:dependsOn ml:AIModel))
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:dependsOn ml:NeuralNetwork))
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:dependsOn ml:TransformerArchitecture))
Capability Relationships
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:enables ml:InferenceServing))
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:enables ml:ModelServing))
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:enables ml:AIInference))
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:enables ml:GenerativeAI))
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:enables ml:OnDeviceAI))
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:enables ml:AutonomousAgents))
Implementation Relationships
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:implements ml:Quantisation))
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:implements ml:TensorParallelism))
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:implements ml:FlashAttention))
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:implements ml:SpeculativeDecoding))
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:uses ml:ONNX))
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:uses ml:GPUAcceleration))
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:uses ml:DistributedSystems))
Reduction Relationships
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:reducesTo ml:InferenceServing))
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:partOf ml:AIInference))
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:partOf ml:MLOps))
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:contrastsWith ml:ModelTraining))
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:contrastsWith ml:FineTuning))
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:relatedTo ml:ModelDeployment))
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:relatedTo ml:CloudComputing))
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:relatedTo ml:AIGovernance))
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:supports ml:LargeLanguageModels))
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:supports ml:RetrievalAugmentedGeneration))
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:supports ml:ComputerVision))
SubClassOf(ml:AIModelInferenceEngine
ObjectSomeValuesFrom(ml:supports ml:NaturalLanguageProcessing))
About
An AI Model Inference Engine is the critical software runtime that bridges the gap between a trained AI Model artefact and live production usage. Whereas AI Model Development produces the frozen computation graph and Model Weights, the inference engine is responsible for executing that graph repeatedly, efficiently, and at scale against a continuous stream of new inputs. The separation between training-time frameworks and inference-time engines reflects a fundamental engineering asymmetry: training requires full gradient computation, dynamic graph construction, and flexibility to experiment with novel architectures, while inference requires deterministic execution, maximum hardware utilisation, low Latency, and predictable memory footprint across thousands of concurrent requests. The economics reinforce this separation: training a frontier Large Language Models costs tens to hundreds of millions of dollars and occurs once or infrequently; serving the resulting model to millions of users happens continuously at a cost that must remain commercially viable on a per-query basis, typically measured in cost per million tokens (CPM). This economic asymmetry has driven a billion-dollar ecosystem of specialised inference optimisation tooling distinct from the training framework ecosystem.
The concept of a dedicated inference engine emerged in the deep learning era as models became too large and computationally demanding to run efficiently on general-purpose training frameworks. Early ONNX (Open Neural Network Exchange) adoption in 2017–2019 established the principle of training-inference separation: train in PyTorch or TensorFlow, export to an interchange format, and run in an optimised engine. NVIDIA TensorRT formalised this pattern for GPU-accelerated inference, compiling ONNX models into hardware-specific kernels that fuse operations, eliminate redundant memory transfers, and apply precision reduction. TensorRT’s profile-based layer fusion evaluates hundreds of candidate kernel implementations for each operation at build time, selecting the fastest for the target GPU and batch configuration. The compilation process (typically 1–20 minutes) front-loads optimisation overhead to build time rather than request-serve time, enabling consistent low-latency serving. The Large Language Models era (2020–present) forced a qualitative evolution: autoregressive Transformer Architecture inference has fundamentally different bottlenecks from the convolutional or feed-forward models TensorRT was originally designed for — specifically, it is memory-bandwidth-bound during the decode phase rather than compute-bound, requires dynamic memory allocation for variable-length sequence KV Caches, and benefits enormously from Continuous Batching rather than static batching. This motivated purpose-built LLM inference engines: Hugging Face TGI (2022), vLLM with PagedAttention (2023), SGLang with RadixAttention (2024), and TensorRT-LLM as NVIDIA’s LLM-specific compiled engine.
The inference engine landscape in 2026 is stratified by deployment context. In data-centre settings, TensorRT-LLM delivers 15–30% higher throughput than vLLM on NVIDIA H100 GPUs for large models (70B+) through ahead-of-time compilation that produces hardware-specific kernels optimised for the specific model graph, weight layout, and batch configuration — at the cost of a compilation step that must be repeated for each new model variant or batch size. SGLang achieves approximately 29% higher throughput than vLLM on smaller models (7B–8B) on H100s and has the lowest tail latency (P95 time-to-first-token) at all tested concurrency levels, making it the preferred choice for latency-sensitive interactive deployments; SGLang’s RadixAttention extends Prefix Caching to tree-structured prompt prefix sharing, enabling cache hit rates of 60–80% in multi-turn conversation and RAG workloads. vLLM remains the most widely adopted open-source LLM inference engine due to its ease of use, broad model support (200+ model architectures), and active community, with its PagedAttention KV cache management eliminating memory fragmentation and enabling effective GPU utilisation across diverse batch compositions; vLLM’s continuous batching maintains 70–90% GPU utilisation under typical API traffic patterns. On edge and consumer hardware, llama.cpp provides CPU-first inference with BLAS-accelerated matrix operations and INT4/INT8 quantised weight loading in the GGUF format, making 7B–70B models accessible on laptop CPUs at 5–15 tokens/sec and consumer GPUs (RTX 4090) at 40–80 tokens/sec for 7B models; Ollama builds a user-friendly model management layer around llama.cpp with automatic model download and multi-model hot-switching; Apple MLX provides a NumPy-like array framework optimised for Apple Silicon (M-series) unified memory architecture, enabling competitive throughput for 7B models on M3 Max (32 GB) without the overhead of separate GPU memory management; and Meta ExecuTorch enables deployment of PyTorch models on Android and iOS via Qualcomm QNN, Core ML, and Vulkan delegates, with sub-20ms latency for Computer Vision and small language model (1B–3B) inference on flagship 2025–2026 mobile SoCs.
Performance Metrics and Benchmarking
The performance of an AI Model Inference Engine is characterised across multiple dimensions that correspond to different production requirements:
Latency Metrics (critical for interactive, user-facing applications):
-
Time-to-First-Token (TTFT): the wall-clock time from request submission to the first output token being returned. TTFT is dominated by the prefill phase (processing the full input prompt) and is the primary user-perceived latency metric for chat interfaces. At high concurrency, TTFT can increase significantly as requests queue for prefill compute. Target SLOs for interactive chat: TTFT < 500ms at P99.
-
Inter-Token Latency (ITL): the time between successive output tokens during the decode phase. ITL is memory-bandwidth-bound and approximately constant at ~30–80ms per token for 70B models on H100 GPUs in non-speculative decoding mode. Speculative decoding reduces effective ITL by 2–3.6×.
-
End-to-End Latency: TTFT + (output_length × ITL). For a 1000-token response at 50ms/token ITL, end-to-end latency is 50 seconds at sustained generation — typically broken by streaming delivery (SSE or WebSocket) that delivers tokens as they are generated.
Throughput Metrics (critical for batch workloads and API providers):
-
Tokens per Second (TPS): the total number of tokens generated per second across all concurrent requests. TPS scales with batch size up to the point where GPU memory is exhausted by KV caches.
-
Requests per Second (RPS): the number of complete requests handled per second. RPS depends on both TPS and average output length.
-
GPU Memory Utilisation (MFU — Memory Footprint Utilisation): the fraction of GPU HBM consumed by Model Weights plus active KV caches. High MFU indicates efficient use of available hardware capacity.
Cost Metrics (critical for commercial viability):
-
Cost per Million Tokens (CPM): the total infrastructure cost (GPU rental + power + networking) per million tokens generated. Reducing CPM is the primary engineering objective for large-scale inference providers.
-
GPU-Hours per Request: total GPU utilisation time per request, accounting for both prefill compute and decode memory bandwidth.
-
H100 GPU rental costs ~20–28/M tokens before any margin — illustrating the economic pressure to maximise TPS.
Quality Metrics (critical for validating that optimisations preserve model accuracy):
-
Accuracy on standard Benchmark Datasets (MMLU, HumanEval, MATH, MT-Bench) measures the impact of Quantisation, pruning, or speculative decoding on model capability. Well-implemented INT4 Quantisation via AWQ or GPTQ typically degrades MMLU accuracy by <1% absolute.
-
Acceptance rate in speculative decoding: the fraction of draft tokens accepted by the target model. High acceptance rates (>75%) are required for speculative decoding to yield net throughput improvements.
-
Sampling faithfulness: the statistical distance between token distributions under optimised engines vs. reference implementations. Deterministic decoding (greedy, beam search) must produce bit-identical results; stochastic sampling (temperature > 0) must preserve the correct output distribution.
MLPerf Inference Benchmarks (industry-standard competitive measurement):
-
MLCommons MLPerf Inference v5.0 (2025) defines measurement rules for data-centre (Offline, Server) and edge (SingleStream, MultiStream) scenarios across model families including Llama 2 70B, GPT-J 6B, BERT 99%, ResNet-50, and SSD-ResNet34.
-
MLPerf Server scenario measures QPS at P99 latency SLO (e.g., P99 TTFT ≤ 2s, P99 ITL ≤ 200ms for LLM); Offline scenario measures peak throughput without latency constraints.
-
NVIDIA H100 SXM (8-GPU) achieves approximately 3,000–4,500 output tokens/sec in MLPerf Inference Server Llama 2 70B with TensorRT-LLM, depending on precision and batch configuration.
Components / Architecture
The AI Model Inference Engine comprises the following architectural layers and subsystems:
-
Computation Graph Execution Engine: The core component that accepts an input tensor (tokenised text, image embeddings, audio frames), loads the model graph, and schedules operations — matrix multiplications, attention layers, activation functions, layer normalisations — onto available compute units. For compiled engines (TensorRT-LLM, Core ML), the graph is fused and compiled ahead-of-time into hardware-specific kernels; for interpreted engines (vLLM, llama.cpp), graph execution is dynamic at the operation level but static at the layer level.
-
KV Cache Manager: The most critical memory management subsystem for autoregressive Transformer Architecture inference. At each decode step, the attention mechanism must attend over all previously generated tokens’ key and value tensors. The KV cache stores these tensors in Hardware Accelerator memory (GPU HBM) to avoid recomputation. Memory footprint scales as
2 × num_layers × num_heads × head_dim × seq_len × batch_size × precision_bytes— for a 70B parameter model with 80 attention layers and 8K context, a single sequence requires approximately 8 GB at FP16, making KV cache management the dominant memory consumer at inference time. PagedAttention (vLLM) divides the KV cache into fixed-size pages (analogous to OS virtual memory paging), eliminating fragmentation and enabling efficient memory sharing across requests with common prefixes. KV cache quantisation to INT8 or FP8 halves memory requirements with negligible quality impact. -
Continuous Batching Scheduler: Dynamic in-flight batching that inserts new requests into a running batch between decode steps rather than waiting for all requests to complete (static batching). Since autoregressive LLMs generate one token per forward pass, the batch composition can change at every step — short requests complete and exit the batch, new requests join at sequence start. This maintains near-constant GPU utilisation regardless of sequence length distribution, achieving 5–10× higher throughput than static batching for interactive workloads.
-
Speculative Decoding Module: An optional throughput-acceleration system in which a small, fast draft model (e.g., a 7B model drafting for a 70B target) auto-regressively proposes K candidate tokens, which the large target model then verifies in a single parallel forward pass — accepting tokens where the draft model’s distribution matches the target’s, and rejecting at the first divergence. When acceptance rates are high (typically 70–85% for well-matched draft/target pairs), speculative decoding yields 2–3.6× throughput improvement with provably identical output distribution to the target model alone. NVIDIA demonstrated 3.6× throughput on H200 GPUs using this technique. By mid-2026, speculative decoding has moved from experimental optimisation to standard practice, with native support in vLLM, TensorRT-LLM, and SGLang.
-
Operator Fusion and Kernel Compilation: Combining multiple sequentially dependent operations (e.g., linear projection → activation → normalisation) into a single GPU kernel invocation, eliminating intermediate tensor materialisation and the associated HBM read/write round-trips. Flash Attention (Dao et al. 2022, 2023) is the canonical example: it tiles the attention computation to stay within GPU SRAM (shared memory), avoiding materialisation of the O(seq_len²) attention score matrix in HBM and achieving both 2–4× speed improvement and 5–10× memory reduction for long-context inference. Custom CUDA and Triton kernels are used for fused QKV projection, RMSNorm-plus-RoPE, and SwiGLU activation sequences.
-
Quantisation Engine: Applies precision reduction to Model Weights and optionally activations to reduce memory footprint and increase arithmetic throughput. Supported formats in leading engines as of 2026 include INT8 (W8A16 weight-only, W8A8 weight-and-activation), INT4 (W4A16, AWQ, GPTQ), FP8 (on NVIDIA Ampere and newer; Hopper FP8 training-aware), FP4 (on NVIDIA Blackwell B100/B200; TensorRT-LLM 0.19+), and BF16/FP16 (baseline). GPTQ and AWQ (Activation-aware Weight Quantisation) are the most widely adopted PTQ (post-training quantisation) methods for LLMs, achieving INT4 compression with typically <1% accuracy degradation on standard benchmarks. SmoothQuant enables INT8 weight-and-activation quantisation by migrating quantisation difficulty from activations to weights through mathematically equivalent channel-wise scaling.
-
Tensor Parallelism and Distributed Systems Coordination: For models whose Model Weights exceed single-GPU memory capacity — typically 70B+ parameters at FP16 require >140 GB VRAM, exceeding a single H100’s 80 GB — tensor parallelism shards individual weight matrices across N GPUs along rows or columns, with each GPU computing a partial matrix product and results reduced via AllReduce across NVLink or InfiniBand. Pipeline parallelism assigns consecutive model layers to different GPUs, streaming micro-batches through to overlap computation and communication. Tensor parallelism degree typically 4–8 for 70B models on H100s; 8–16 for 175B–405B models.
-
Prefix Caching: Reuse of KV cache computations for repeated prompt prefixes across requests. In Retrieval-Augmented Generation workloads, system prompts, few-shot examples, or retrieved context documents appear in every request; prefix caching avoids recomputing these during the prefill phase. Provider-level prompt caching is deployed commercially by Anthropic, OpenAI, and Google Gemini APIs. Radix attention (SGLang) generalises prefix caching to tree-structured cache sharing across arbitrarily branching prompt prefixes, improving cache hit rates for complex agentic workflows.
-
ONNX Interoperability Layer: Open Neural Network Exchange format provides a hardware-agnostic representation of trained computation graphs that enables portability between training frameworks and inference engines. ONNX Runtime (Microsoft) is the cross-platform inference engine supporting CPU, GPU (CUDA, TensorRT), NPU (DirectML, QNN), and mobile targets. The ONNX ecosystem is the primary interoperability mechanism for Computer Vision models deployed across heterogeneous hardware, from cloud GPUs to embedded NPUs, making it central to Edge Computing and On-Device AI deployment pipelines.
Use Cases / Major Families
AI Model Inference Engine implementations span several major deployment families:
-
Cloud LLM Inference Engines: vLLM (meta-engine for open-weight models: PagedAttention, continuous batching, prefix caching, speculative decoding, 200+ supported model families), SGLang (RadixAttention, structured output generation, lowest P95 latency), TensorRT-LLM (NVIDIA’s compiled engine for maximum throughput on NVIDIA hardware), and Text Generation Inference — TGI (Hugging Face, production-grade, safetensors format, tensor parallelism). These power commercial inference APIs: AWS Bedrock, Google Vertex AI, Azure AI Studio, Groq (custom LPU hardware with TensorRT-LLM-equivalent compilation), Cerebras (wafer-scale engines), and Together AI.
-
General-Purpose Deep Learning Inference Runtimes: ONNX Runtime (cross-platform, multi-accelerator, production-grade for Computer Vision and structured ML), NVIDIA TensorRT (compiled, highest-throughput for NVIDIA GPU-deployed convolutional and vision transformer models), OpenVINO (Intel hardware — CPU, GPU, NPU, FPGA), Core ML / MLX (Apple Silicon, unified memory, Neural Engine acceleration), and TensorFlow Lite / LiteRT (Google’s mobile and embedded inference runtime).
-
Edge and On-Device Runtimes: llama.cpp (C++ LLM inference on CPU and GPU, GGUF weight format, INT4/INT8 quantisation, runs Llama 3 on laptop CPUs at ~5–15 tokens/sec), Ollama (llama.cpp wrapper with model management), Apple MLX (NumPy API over Apple Silicon GPU/Neural Engine, FP16/INT4, competitive with cloud inference for 7B models on M3 Max), Meta ExecuTorch (PyTorch-to-mobile deployment, Qualcomm QNN and Core ML delegates for NPU acceleration, sub-20ms latency on mobile in 2026), Nexa AI SDK (multimodal on-device inference, ARM NPU support), and MediaPipe (Google, real-time Computer Vision and audio on mobile/edge).
-
Specialised Inference Engines: Whisper.cpp (ASR inference), Stable Diffusion inference servers (AUTOMATIC1111, ComfyUI, diffusers with CUDA graph optimisation), AlphaFold 3 inference (protein structure prediction on GPU clusters), and NVIDIA Triton Inference Server (multi-model, multi-framework serving orchestration for enterprise deployments).
-
Agentic Multi-Step Inference: Inference engines increasingly serve as the substrate for Autonomous Agents that make multiple sequential inference calls with tool-calling between steps. SGLang’s structured output generation and RadixAttention are specifically optimised for agentic patterns; caching across agent reasoning steps and batching of tool-call results are active engineering challenges in 2026.
Hardware Accelerator Landscape
The inference engine’s performance is fundamentally determined by the Hardware Accelerator it targets. Each accelerator family presents different compute-memory trade-offs that engine design must exploit:
NVIDIA GPU Family (Data-Centre):
-
H100 SXM5: 80 GB HBM3 at 3.35 TB/s, 3,958 TFLOPS BF16 Tensor Core, 900 GB/s NVLink 4.0. The dominant data-centre inference accelerator in 2025–2026. TensorRT-LLM and vLLM are both optimised for H100 SXM architecture.
-
H200 SXM5: 141 GB HBM3e at 4.8 TB/s (43% more memory bandwidth than H100), enabling serving of larger models without tensor parallelism (70B at FP16 fits within H200’s 141 GB). 3.6× speculative decoding speedup demonstrated on H200 by NVIDIA.
-
B100/B200 Blackwell (2025–2026 deployment): 192 GB HBM3e at 8 TB/s, FP4 Tensor Cores (5× the FLOPS of H100 at FP8), and NVLink 5.0 at 1.8 TB/s per GPU. TensorRT-LLM FP4 mode on B200 enables serving 405B+ parameter models on 8-GPU nodes with sub-100ms TTFT for interactive workloads.
-
RTX 4090 (Consumer): 24 GB GDDR6X at 1 TB/s. Widely used for local inference via vLLM with INT4 quantised models; a single RTX 4090 serves 7B models at ~80 tokens/sec and 70B models at ~10 tokens/sec with GPTQ INT4 quantisation.
Google TPU Family:
-
TPU v5p (2024): 95 GB HBM2e per chip, interconnected via a high-bandwidth 2D torus network at 4,800 GB/s/chip; optimised for JAX/XLA model compilation. Used internally for Gemini inference.
-
TPU v6 (Trillium, 2025): 5× FLOPS improvement over TPU v5e, targeting both training and inference workloads. Available on Google Cloud Vertex AI for inference serving.
AMD Instinct GPU Family:
-
MI300X (2024): 192 GB HBM3 at 5.3 TB/s per GPU, 3× H100’s memory capacity. ROCm software stack enables vLLM and TensorRT-LLM compatibility. Used by Microsoft Azure for large-scale open-weight model inference, and by Meta for Llama 3 inference.
-
MI325X (2025): 256 GB HBM3e at 6 TB/s, positioning AMD competitively for 70B+ model inference without tensor parallelism.
Edge and Mobile NPUs:
-
Qualcomm Snapdragon 8 Gen 3/Gen 4: Hexagon NPU, 45+ TOPS (tera-operations per second) on-device AI performance, INT4/INT8 inference for LLMs up to 7B parameters. ExecuTorch + QNN delegate enables sub-20ms Computer Vision model inference in 2026.
-
Apple Neural Engine (ANE): integrated in A18 Pro (iPhone 16 Pro) and M4/M4 Pro. 38 TOPS on A18 Pro; Apple MLX framework provides the primary inference path for on-device LLM applications on iOS and macOS, supporting FP16 and INT4 precisions with Metal compute shaders.
-
ARM Cortex-A NPU / Ethos series: ARM Ethos-N78 integrated NPU in Cortex-A78 SoCs; Kleidi AI libraries from ARM provide optimised INT4/INT8 inference kernels targeting the ~90% of smartphones using ARM Cortex-A CPUs.
-
NVIDIA Jetson Orin: 275 TOPS on-device inference platform for edge robotics, autonomous vehicles, and industrial automation; runs TensorRT-compiled models for Computer Vision and small LLMs.
Specialised AI Accelerators:
-
Groq LPU (Language Processing Unit): Deterministic SRAM-based architecture eliminating HBM bandwidth bottleneck; achieves 500+ tokens/second for Llama 3 70B with Groq’s proprietary compiler. No on-chip weight loading from HBM means weight tensors are fully resident in fast SRAM during inference.
-
Cerebras WSE-3 (Wafer-Scale Engine): single-chip 4TB/s on-chip bandwidth, 900,000 AI cores; eliminates inter-chip communication overhead for models that fit on the wafer (up to ~200B parameters). Used for ultra-fast single-request inference with sub-10ms TTFT.
-
Graphcore IPU-M2000: 1.47 PFLOPS, 900 MB SRAM on-chip, BSP (Bulk Synchronous Parallel) execution model; suited for inference with small batch sizes where HBM bandwidth is insufficient.
Optimisation Strategies and Engineering Patterns
The AI inference engineer’s toolkit for maximising throughput and minimising latency at target quality:
Memory Bandwidth Optimisation (primary bottleneck in LLM decode phase):
-
Weight-only quantisation (W4A16, W8A16): quantise stored Model Weights to INT4/INT8 while computing activations in FP16. This reduces the bytes transferred from HBM per token generated, directly improving decode throughput. AWQ INT4 typically achieves 3.5–4× weight loading bandwidth reduction vs. FP16.
-
Grouped Query Attention (GQA): share key-value projection matrices across multiple query heads (e.g., 8 query heads share 1 KV head in Llama 3), reducing KV Cache memory footprint by 4–8× vs. Multi-Head Attention (MHA) without GQA. GQA is now standard in all frontier LLMs (Llama 3, Mistral, Qwen 2.5, Gemma 2).
-
Speculative Decoding: amortises memory bandwidth by verifying K draft tokens in a single forward pass rather than generating each individually, reducing effective memory bandwidth per output token by 2–3.6×.
-
KV Cache quantisation (FP8/INT8 KV): applies precision reduction to stored KV tensors, halving KV Cache memory at the cost of ~0.3% MMLU accuracy. Supported in TensorRT-LLM and vLLM as of 2025.
Compute Throughput Optimisation (primary bottleneck in prefill phase):
-
Continuous Batching: maximises GPU utilisation by keeping the batch full across decode steps as requests complete and new ones join.
-
Flash Attention and FlashAttention-2/3: tiled attention computation within SRAM, eliminating O(seq_len²) HBM accesses for attention score materialisation. FlashAttention-3 on H100 uses the Tensor Memory Accelerator (TMA) for asynchronous SRAM loading, achieving 740 TFLOPS BF16 (75% of H100 peak) for attention computation.
-
Chunked prefill (Sarathi-Serve, 2024): splits large prefill requests into fixed-size chunks interleaved with decode steps, preventing prefill from blocking the decode batch and reducing TTFT variability at high concurrency.
-
CUDA graph capture: records the sequence of GPU kernel launches for a fixed batch size and replays them from a command buffer, eliminating CPU-GPU scheduling overhead (~5–10% throughput improvement for small batch sizes).
Compilation and Kernel Fusion:
-
torch.compilewith Inductor backend: JIT compilation of PyTorch model graphs into fused TorchInductor kernels for CPU/CUDA; reduces kernel launch overhead and applies loop fusion and tiling. 30–50% throughput improvement on common inference patterns. -
TensorRT compilation pipeline: ONNX graph → TensorRT builder → profile-optimised TRT engine; layer fusion identifies fusible operator sequences and replaces them with single kernel calls; INT8/FP8/FP4 calibration integrated into build step.
-
Triton kernel development: OpenAI Triton DSL enables writing custom GPU kernels at a higher level than CUDA with competitive performance; used for fused attention variants, quantised matmuls, and custom activation functions in vLLM and SGLang.
Structured Output Generation:
-
Grammar-constrained decoding: inference engines apply finite-state machine constraints over the token vocabulary at each decode step, forcing outputs to conform to a JSON schema, regular expression, or context-free grammar. LMQL, Outlines (Dottxt AI), and SGLang’s constrained generation implement this.
-
Structured output generation is essential for reliable agentic Autonomous Agents tool calling, API integration, and downstream parser compatibility, enabling inference engines to produce JSON, SQL, or code that can be executed without post-processing error handling.
Academic Context
The theoretical and algorithmic foundations of AI inference engines draw from computer architecture, compiler design, numerical linear algebra, and machine learning research. Key contributions include:
Attention mechanism efficiency research is foundational: Bahdanau et al. (2015) introduced soft attention for NMT; Vaswani et al. (2017) Transformer paper established the architecture that dominates LLM inference; Kitaev et al. (2020) Reformer introduced LSH attention for long sequences; Child et al. (2019) Sparse Transformer proposed sparse attention patterns. Dao et al. (2022) Flash Attention and Dao (2023) FlashAttention-2 are the most impactful system-level contributions to LLM inference, with FlashAttention-3 (2024) extending to H100’s tensor memory accelerator (TMA) and FP8. Kwon et al. (2023) introduced PagedAttention in the vLLM paper, the foundational contribution enabling efficient KV cache management for concurrent request serving. Leviathan et al. (2023) and Chen et al. (2023) independently proposed speculative decoding, formalising the theoretical framework for draft-verify inference acceleration. The ONNX specification (Bai et al., 2019) standardised cross-framework model interchange. NVIDIA TensorRT’s layer fusion and kernel auto-tuning build on classical compiler loop optimisation and polyhedral transformation techniques. Quantisation research for Neural Networks spans Gupta et al. (2015, stochastic rounding), Jacob et al. (2018, quantisation-aware training for integer arithmetic), Frantar et al. (2022, GPTQ), and Lin et al. (2023, AWQ). Agrawal et al. (2024) Sarathi-Serve addressed prefill-decode interference through chunked prefill scheduling. Zhong et al. (2024) DistServe proposed disaggregated prefill-decode as a systems architecture for maximising cluster-level goodput. The FlashInfer library (2025) provides a highly modular attention kernel library for LLM inference serving, enabling rapid integration of new attention variants (sliding window, cross-attention, MLA as in DeepSeek-V2) into inference engines without full reimplementation.
Security and Safety Considerations for Inference Engines
Inference engines introduce a distinct security and safety surface that differs from training-time concerns:
-
Prompt Injection Attacks: Adversarial instructions embedded in retrieved documents, tool outputs, or user-controlled inputs can override the system prompt’s intended behaviour, causing the model to exfiltrate data, bypass safety guardrails, or execute unintended tool calls. Inference-time defences include input sanitisation (removing known injection patterns before tokenisation), instruction hierarchy enforcement (OpenAI’s meta-prompt architecture that explicitly scopes user-provided content), and output monitoring (post-generation classifiers that detect instruction-following of injected commands rather than the original system prompt). SGLang’s constrained generation can prevent certain injection patterns by enforcing output schema constraints that rule out exfiltration patterns.
-
Model Extraction Attacks: Repeated queries to a black-box inference API can enable adversaries to approximate model weights or decision boundaries through model stealing attacks (Tramèr et al. 2016). Defences employed by inference providers include: output perturbation (adding small stochastic noise to logit distributions), rate limiting (throttling requests per API key to prevent systematic extraction), watermarking (statistically embedded signatures in generated token distributions detectable by the model owner), and query pattern detection (anomaly detection on request sequences consistent with systematic probing). The inference engine’s role in this defence lies in implementing output sampling with controlled randomness and query logging for forensic analysis.
-
Data Residency and Privacy: Inference workloads processing sensitive data — patient health records, legal documents, financial transactions — face GDPR, HIPAA, and sector-specific regulatory constraints on where computation may occur and how inference logs may be retained. On-premise inference engines (self-hosted vLLM, Ollama) and On-Device AI runtimes address data residency by ensuring inputs never leave controlled infrastructure. Inference engines in regulated deployments must implement: structured log retention policies (inference inputs/outputs may constitute personal data), differential privacy output mechanisms (adding noise to prevent membership inference from outputs), and audit trails for AI Governance compliance.
-
Reproducibility and Determinism: GPU floating-point non-determinism (from parallel reduction order sensitivity in CUDA kernels) and sampling randomness (temperature, top-p nucleus sampling) mean that identical inputs can produce different outputs across inference runs or across different GPU models. Deterministic inference — required for testing, debugging, and some regulatory audit scenarios — requires fixed random seeds, deterministic CUDA kernels (CUDA
--use-fast-mathdisabled, NCCL deterministic mode), and greedy (argmax) decoding. Deterministic mode typically reduces throughput by 5–15% due to disabling non-deterministic but faster reduction algorithms. -
Adversarial Robustness of Inference Outputs: While adversarial robustness is primarily a training-time concern, inference engines can apply input preprocessing defences including: random smoothing (perturbing inputs before inference to smooth out adversarial perturbations), input transformation defences (JPEG compression, bit-depth reduction for image inputs), and certified ensemble methods. For LLMs, inference-time defences against jailbreaks include output filters (Llama Guard, ShieldGemma, PromptGuard) that classify inference outputs as safe/unsafe before delivery to the user.
Current Landscape (2026)
The 2026 AI Model Inference Engine landscape is defined by three concurrent trends. First, the LLM inference engine market has consolidated around three primary open-source engines — vLLM, SGLang, and TensorRT-LLM — with the competitive differentiation shifting from feature parity to latency, memory efficiency, and model support breadth. H100 GPU benchmarks published in 2026 show TensorRT-LLM achieving 15–30% higher throughput than vLLM on 70B+ models through compiled kernel optimisation, SGLang achieving ~29% higher throughput on 7B–8B models with lower tail latency, and all three engines supporting speculative decoding (2–3.6× speedup, now standard), FP8 precision (NVIDIA Ampere+), and Prefix Caching. The FlashInfer-Bench project (arXiv:2601.00227, 2025) is developing standardised benchmarking methodology for LLM inference engines, addressing the reproducibility gap in published throughput comparisons. The inference-as-a-service market has expanded significantly: AWS Bedrock, Google Vertex AI, Azure AI Studio, and Hugging Face Inference Endpoints now provide managed inference endpoints abstracting all engine selection, scaling, and optimisation from the application developer, with pay-per-token pricing models enabling cost-efficient access without infrastructure management. Specialised inference providers — Groq (LPU-based, 500+ tokens/sec for Llama 3 70B), Cerebras (WSE-3, wafer-scale inference), Together AI, Fireworks AI, and Lepton AI — compete on throughput, latency, and open-weight model breadth. Second, on-device inference has achieved production-quality results: by mid-2026, sub-20ms latency for Computer Vision models on flagship Android (Snapdragon 8 Gen 4) and iOS (A18 Pro) devices is standard via NPU-accelerated ExecuTorch and Core ML deployments, enabling privacy-preserving, offline-capable On-Device AI applications. Apple MLX delivers competitive throughput for 7B LLMs on M3 Max (32 GB unified memory) without quantisation (~25–30 tokens/sec), while Ollama with llama.cpp achieves 40–80 tokens/sec for 7B INT4 models on consumer RTX 4090 GPUs. Third, NVIDIA’s Blackwell architecture (B100/B200, delivered at scale in 2025–2026) introduces FP4 tensor cores (4× dense FLOPS vs. H100), 8 TB/s HBM3e bandwidth (2.4× H100), and NVLink 5.0 at 1.8 TB/s per GPU, fundamentally changing the memory-bandwidth envelope for LLM inference and enabling TensorRT-LLM FP4 inference with negligible accuracy loss. The GB200 NVL72 — 72 B200 GPUs connected via NVLink with 1.4 TB/s inter-GPU bandwidth — enables serving 405B+ parameter models with sub-50ms TTFT, meeting interactive latency requirements for the largest open-weight models without tensor parallelism across InfiniBand clusters.
UK Context
The UK has a substantive AI inference engineering and research presence, spanning hardware architecture, systems research, and applied deployment:
Industry Anchors:
-
ARM Holdings (Cambridge, FTSE 100 following September 2023 Nasdaq IPO; primary shareholder SoftBank) is the world’s dominant processor architecture licensor, with ARM Cortex-A and Cortex-X cores used in >99% of smartphones and the majority of IoT devices. ARM’s Kleidi AI libraries provide highly optimised inference kernels for INT4/INT8 quantised LLMs on ARM Cortex-A CPUs, targeting the >90% of mobile SoCs that use ARM architecture for On-Device AI workloads. ARM’s Ethos NPU IP (Ethos-N78, Ethos-U85) provides dedicated neural network acceleration integrated alongside CPU cores, targeting 10–100 TOPS for edge inference applications. ARM’s Cambridge Research Centre contributes to ONNX Runtime optimisation for ARM architectures and participates in mobile NPU inference standardisation alongside Qualcomm, Apple, and MediaTek.
-
Google DeepMind (London, ~2,500 staff) has made foundational contributions to inference efficiency research: speculative decoding (Leviathan et al. 2023, in joint Google Brain/DeepMind authorship), mixture-of-experts inference routing (Switch Transformer, Fedus et al. 2021), Flash Attention-style attention reformulations, and the AlphaFold inference serving infrastructure for high-throughput protein structure prediction. DeepMind’s Gemini Ultra/Pro/Flash series serves as inference benchmarks for the UK’s national AI infrastructure assessment.
-
Graphcore (Bristol, UK; founded 2016 by Nigel Toon and Simon Knowles): developed the Intelligence Processing Unit (IPU), a MIMD (Multiple Instruction, Multiple Data) processor with 900 MB on-chip SRAM (MK2 GC200) per chip, enabling inference without the HBM memory bandwidth bottleneck that constrains GPU-based engines. The IPU’s Bulk Synchronous Parallel (BSP) execution model eliminates the need for cache coherency and enables highly deterministic latency profiles suited for latency-SLO-sensitive inference workloads. Despite financial difficulties in 2024, Graphcore’s IPU technology was acquired by SoftBank in 2024 as part of ARM integration planning, potentially influencing future ARM IP roadmap for edge inference acceleration.
-
Wayve (London): applies neural inference to end-to-end autonomous driving, with deployed inference on NVIDIA Orin-based vehicle computers running proprietary neural driving policy models at 25–60Hz with hard real-time requirements — demanding inference engine reliability guarantees that standard cloud-serving engines are not designed for.
University Research Groups:
-
University of Edinburgh’s EPCC (Edinburgh Parallel Computing Centre): provides ARCHER2 (UK national supercomputer, 748,544 cores, 34 PFLOPS) and Cirrus GPU cluster infrastructure used for large-scale batch AI Inference benchmarking, LLM Model Evaluation at scale, and inference serving research. EPCC collaborates with the Software Sustainability Institute on reproducibility standards for ML inference benchmarks.
-
Imperial College London’s HPC Service and Machine Learning group contribute to distributed inference systems research; the Distributed Algorithms Group investigates fault-tolerant inference serving under partial accelerator failures, relevant for production reliability engineering.
-
UCL’s NVIDIA partnership (2025) includes access to DGX H100 cluster infrastructure for inference benchmarking and optimisation research. UCL’s AI Centre maintains a production inference platform serving the UCL research community and NHS partner organisations with open-weight LLM inference under data-residency constraints.
-
University of Cambridge’s Machine Intelligence Laboratory (MIL): contributes to Bayesian inference and probabilistic ML — including uncertainty quantification at inference time — which is increasingly relevant for deploying AI Model inference in safety-critical healthcare and scientific applications where calibrated confidence estimates are required.
Northern English Industrial Context:
-
Sheffield’s Advanced Manufacturing Research Centre (AMRC) and Digital Manufacturing on a Shoestring programme apply inference engine deployment to low-cost industrial IoT devices, using ONNX Runtime and TensorFlow Lite on Raspberry Pi and ARM Cortex-M devices for real-time manufacturing quality control inference within constrained edge environments.
-
Leeds Teaching Hospitals NHS Trust (in collaboration with University of Leeds AI Institute): deploys inference engines for medical imaging diagnostics — chest X-ray pneumonia detection, MRI brain tumour segmentation — on-premises with NHS data-residency constraints, using NVIDIA Triton Inference Server with TensorRT-compiled models.
-
Newcastle’s Digital Health team (in partnership with NHS North East and North Cumbria ICS) runs clinical NLP inference using locally hosted Mistral 7B and Llama 3 8B models via vLLM, processing GP discharge summaries and clinical notes with GDPR-compliant on-premises inference to extract structured clinical data.
-
Manchester Digital City Region initiative: supports mid-market UK enterprises deploying Ollama-based local inference on business workstations for document analysis and customer service automation, avoiding per-token cloud API costs and data residency concerns, with particular uptake in the financial services sector concentrated in Manchester’s Spinningfields district.
Future Directions (2026–2030)
-
Disaggregated Prefill-Decode Architectures: Physical separation of the compute-bound prefill phase and the memory-bandwidth-bound decode phase onto different, purpose-optimised accelerator pools. Prefill nodes use high-FLOPS accelerators (H100/H200 compute-optimised instances), while decode nodes use high-memory-bandwidth accelerators (H200/MI300X memory-optimised instances), each running at optimal utilisation rather than underperforming in the mixed-workload regime. DistServe (Zhong et al. 2024, OSDI) demonstrated goodput improvements of 2–4× over collocated prefill-decode serving by eliminating the interference between compute-intensive prefill and memory-intensive decode. SplitWise (Patel et al. 2024) extended this to heterogeneous accelerator clusters. Production adoption by 2027–2028 is anticipated as cluster sizes grow to support the additional routing overhead.
-
FP4 and Sub-4-Bit Quantisation: NVIDIA Blackwell’s FP4 tensor cores (delivering 5× H100’s INT8 FLOPS) and TensorRT-LLM 0.19+ FP4 mode enable 405B+ parameter models within 8-GPU B200 NVL8 nodes (192 GB × 8 = 1.5 TB GPU memory). Sub-4-bit research — BitNet b1.58 (Ma et al. 2024, arXiv:2402.17764) theorising 1.58-bit weights equivalent to ternary {-1, 0, 1} values — promises 8× weight memory reduction vs. FP16 for natively-quantised-trained models, with hardware implementations on NVIDIA Rubin (2027) expected to support sub-4-bit arithmetic. Model Evaluation on Humanity’s Last Exam and complex coding Benchmark Datasets will determine whether capability is preserved at extreme quantisation.
-
Neuromorphic and Non-Von Neumann Inference: IBM NorthPole (2023) achieves 22 TOPS/W at INT2 precision on a single chip by placing 256 MB SRAM directly adjacent to compute cores, eliminating off-chip DRAM accesses entirely for models that fit within the 256 MB footprint. Intel Loihi 2 (2021) and its successors implement spiking Neural Network inference at sub-milliwatt power levels for always-on sensor processing applications. As LLM inference remains memory-bandwidth-bound in the decode phase at ~200–500 tokens/second on H100s, architectures that eliminate the off-chip HBM bottleneck by design — Groq LPU, Cerebras WSE-3, Graphcore IPU — may achieve structural efficiency advantages for medium-size model families where the full weight set fits on-chip.
-
Cross-Model Batching and Multi-Tenant Inference: Serving multiple model families from a shared inference engine with dynamic GPU memory allocation, enabling a single H100 cluster to serve simultaneous requests for LLaMA 3, Whisper, and Stable Diffusion XL without static memory partitioning or physical server separation. vLLM v0.7+ prefix-sharing-aware memory management and SGLang’s multi-model support (2026) begin this trajectory; production multi-tenant inference scheduling that dynamically allocates KV cache memory across co-running models based on demand signals is the next step.
-
Inference-Time Alignment and Safety Layers: Integration of safety classifiers (Llama Guard 3, ShieldGemma, PromptGuard), AI Governance compliance checkers, structured output validators, and prompt injection detectors directly into the inference engine execution path — operating on token streams before and after generation with sub-5ms additional latency overhead — rather than as external post-processing services that add 50–200ms round-trip latency. Kernel-fused safety scoring that executes alongside the model’s final linear layer projection is the efficiency target for 2027–2028 implementations.
-
Energy-Efficient Inference Standards: MLCommons’ energy-efficiency benchmark extensions (tokens per joule, inference operations per watt, carbon emissions per million tokens) will supplement throughput metrics, driven by data-centre power constraints. NVIDIA B200 achieves approximately 4× the energy efficiency of H100 at equivalent throughput for FP4 LLM inference. EU AI Act sustainability reporting requirements effective 2026 will mandate energy consumption disclosure for large-scale inference deployments. Microsoft’s Azure commitment to 100% renewable energy for data-centre operations by 2025 and net-zero data-centre water consumption by 2030 are driving power-proportional inference scheduling research.
-
Inference-Time Scaling (Test-Time Compute): The emerging paradigm of reasoning models — DeepSeek-R1, OpenAI o3, Claude 3.7 extended thinking — trades additional inference compute (extended chain-of-thought token generation) for improved accuracy on hard reasoning tasks. Inference engines must efficiently serve these longer-output-sequence workloads, where a single response may generate 2,000–20,000 tokens, fundamentally changing the batch composition and KV Cache memory allocation patterns compared to short-output chat workloads. Efficient reasoning inference requires specialised scheduling that prioritises long-context decode throughput over TTFT minimisation.
Research & Literature
- Vaswani, A., et al. (2017). Attention Is All You Need. NeurIPS, 30.
- Dao, T., Fu, D., Ermon, S., Rudra, A., & Ré, C. (2022). FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness. NeurIPS, 35.
- Dao, T. (2023). FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning. ICLR 2024.
- Shah, J., et al. (2024). FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision. arXiv:2407.08608.
- Kwon, W., et al. (2023). Efficient Memory Management for Large Language Model Serving with PagedAttention. SOSP 2023.
- Leviathan, Y., Kalman, M., & Matias, Y. (2023). Fast Inference from Transformers via Speculative Decoding. ICML 2023.
- Chen, C., et al. (2023). Accelerating Large Language Model Decoding with Speculative Sampling. arXiv:2302.01318.
- Agrawal, A., et al. (2024). Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve. OSDI 2024.
- Sheng, Y., et al. (2024). FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU. ICML 2023.
- Zheng, L., et al. (2023). Efficiently Programming Large Language Models using SGLang. NeurIPS 2024.
- NVIDIA. (2024). TensorRT-LLM: A TensorRT Toolbox for Optimized Large Language Model Inference. NVIDIA Developer Blog.
- Frantar, E., Ashkboos, S., Hoefler, T., & Alistarh, D. (2022). GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers. ICLR 2023.
- Lin, J., et al. (2023). AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration. MLSys 2024.
- Xiao, G., et al. (2022). SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models. ICML 2023.
- Bai, J., et al. (2019). ONNX: Open Neural Network Exchange. https://github.com/onnx/onnx.
- Jacob, B., et al. (2018). Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference. CVPR 2018.
- Micikevicius, P., et al. (2018). Mixed Precision Training. ICLR 2018.
- Ma, S., et al. (2024). The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits. arXiv:2402.17764.
- Patel, P., et al. (2024). Splitwise: Efficient Generative LLM Inference Using Phase Splitting. arXiv:2311.18677.
- Zhong, Y., et al. (2024). DistServe: Disaggregating Prefill and Decoding for Goodput-Optimized Large Language Model Serving. OSDI 2024.
- Shoeybi, M., et al. (2019). Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism. arXiv:1909.08053.
- FlashInfer Team. (2025). FlashInfer-Bench: Building the Virtuous Cycle for AI-driven LLM Systems. arXiv:2601.00227.
- Hugging Face. (2023). Text Generation Inference. https://github.com/huggingface/text-generation-inference.
- vLLM Team. (2025). Inside vLLM: Anatomy of a High-Throughput LLM Inference System. vLLM Blog.
- Yotta Labs. (2026). Best LLM Inference Engines (2026): vLLM, SGLang & TensorRT-LLM Compared. https://www.yottalabs.ai/post/best-llm-inference-engines-in-2026-vllm-tensorrt-llm-tgi-and-sglang-compared.
- Spheron Network. (2026). vLLM vs TensorRT-LLM vs SGLang: H100 Benchmarks (2026). https://www.spheron.network/blog/vllm-vs-tensorrt-llm-vs-sglang-benchmarks/.
- Aleph Zero Labs. (2026). On-Device AI Inference in 2026: Sub-20ms on Android, Real Benchmarks, and When to Go Edge. https://www.alephzerolabs.com/blog/on-device-ai-2026-sub-20ms/.
- MLCommons. (2025). MLPerf Inference v5.0 Results. MLCommons.org.