Low latency is an engineering design property characterising systems in which the elapsed time between an input event and the corresponding system response is minimised to meet real-time interaction requirements. It is a cross-cutting concern spanning network topology, compute placement, operating-system scheduling, memory hierarchy, and hardware design. Achievable thresholds are domain-dependent — sub-100 µs in high-frequency trading, under 20 ms for vestibulo-ocular reflex alignment in extended-reality headsets, and under 150 ms for interactive video — and are reached through a combination of edge computing, kernel-bypass networking, hardware acceleration, and optimised serialisation. As a foundational property of distributed infrastructure, low latency is a prerequisite for real-time AI inference, immersive spatial computing, autonomous robotics, and ultra-reliable industrial control.

Overview

  • Low latency matters because human perception and physical processes impose hard time budgets. The vestibulo-ocular reflex imposes a ~12 ms budget on XR rendering; industrial servo loops need sub-millisecond cycle times; trading algorithms respond to market events faster than human reaction. In all cases, exceeding the budget degrades correctness (stale data acted upon), safety (delayed robot braking), or user experience (motion sickness, conversational disruption).
  • Latency decomposition — the classical view — splits total delay into four additive components:
    • Propagation delay: governed by the speed of light over the physical path; minimised by reducing geographical distance or using Co-Location.
    • Transmission delay: packet size divided by link bandwidth; minimised through efficient serialisation, smaller messages, and higher-bandwidth links.
    • Processing delay: compute time in switches, servers, or sensors; minimised through Hardware Acceleration (GPUs, FPGAs, ASICs) and lock-free algorithms.
    • Queuing delay: congestion-induced waiting; minimised through Quality Of Service scheduling, priority queues, and traffic shaping.
  • Each component must be addressed separately; improvements to one cannot compensate for neglect of another at the microsecond scale.

Key Mechanisms

  • Edge Computing: Deploys compute nodes within the network, minimising propagation delay by processing data near the source. Multi-access edge computing (MEC) embeds compute within 5G base stations for URLLC (Ultra-Reliable Low-Latency Communication) scenarios.
  • Kernel-Bypass Networking: Frameworks such as DPDK (Data Plane Development Kit) and RDMA allow applications to exchange network packets without kernel-space transitions, eliminating OS scheduling jitter and achieving sub-5 µs latencies in data-centre fabrics.
  • FPGA and Custom Silicon: Field-programmable gate arrays and application-specific integrated circuits (ASICs) implement packet parsing, order matching, or signal processing in hardware pipelines with deterministic, nanosecond-scale latency, avoiding software stacks entirely.
  • Real-Time Operating System: Kernels configured with PREEMPT_RT or real-time extensions prevent unbounded preemption by the scheduler, providing bounded worst-case interrupt latency (jitter control).
  • Lock-Free Data Structures: Concurrency without mutual exclusion eliminates priority inversion and contention stalls in multi-threaded hot paths.
  • Content Delivery Network (CDN): Geographically distributes static and semi-static content to reduce round-trip time for end-user requests.
  • Network Slicing: Isolates dedicated virtual network segments with guaranteed bandwidth and latency budgets within shared physical infrastructure, a key feature of 5G core networks.
  • WebRTC: A browser-native protocol stack designed for peer-to-peer real-time audio/video with sub-200 ms end-to-end latency, combining ICE, DTLS-SRTP, and congestion control.
  • Busy-Polling: Threads continuously spin on network queues or memory locations rather than sleeping, trading CPU cycles for predictable, jitter-free detection of incoming events.
  • Speculative Decoding and Batching: In Real-Time AI Inference, techniques such as speculative decoding, continuous batching, and quantisation reduce per-token generation time to match conversational latency budgets.

Applications and Use Cases

  • High-Frequency Trading: Co-location of trading engines at exchange data centres, microwave/millimetre-wave relay links, and FPGA-based order matching exploit every microsecond of advantage. Co-location services are the canonical industrial expression of latency arbitrage.
  • Extended Reality (XR): Immersive headsets and holographic displays require motion-to-photon latency under 20 ms. Apple Vision Pro’s R1 processor, dedicated to sensor fusion and display processing, is an instance of custom silicon solving a latency budget in consumer hardware.
  • Real-Time AI Inference: Conversational AI, autonomous driving perception, and robotics control loops all require inference pipelines that complete within the application’s time budget. Batching, quantisation (INT8, FP8), and model distillation reduce compute latency; edge deployment reduces network latency.
  • Industrial Automation and Robotics: Closed-loop servo control in CNC machines, robot arms, and autonomous vehicles requires sub-millisecond sensor-to-actuator cycles. Real-Time Operating System configurations and EtherCAT fieldbus protocols are standard solutions.
  • Tactile Internet: The concept of networked haptic feedback (remote surgery, teleoperation) requires round-trip latencies under 1 ms over wide areas — a still-unsolved challenge at scale and a long-term motivator for 5G URLLC and photonic networking.
  • Cloud Gaming and Interactive Video: Streaming render output to thin clients requires latency low enough that controller input lag is imperceptible; major providers target under 40 ms glass-to-glass.
  • Distributed Databases and Consensus Protocols: Low-latency storage engines and Raft/Paxos implementations in data centres optimise for single-digit millisecond commit times; RDMA-backed fabrics enable sub-millisecond RPC within a rack.
  • Spatial Computing: Augmented and mixed reality overlays must track physical objects and re-render at display rate; photon latency directly impacts perceived stability of virtual anchors.
  • Autonomous Vehicles: Lidar, radar, and camera fusion pipelines must complete in under 100 ms for safe manoeuvring at highway speeds; dedicated SoCs (Nvidia Drive, Mobileye EyeQ) implement hardware-accelerated perception graphs.
  • Edge AI: Inferring on-device or at a nearby MEC node eliminates cloud round-trip latency, enabling always-on intelligence in constrained network environments.

Standards and Context

  • IETF DiffServ (RFC 2474/2475): Defines Differentiated Services for IP networks, enabling per-hop prioritisation of latency-sensitive traffic classes.
  • IETF IntServ (RFC 1633): Resource Reservation Protocol (RSVP) for end-to-end guaranteed service; precursor to modern QoS frameworks.
  • IEEE 802.1Qbv (Time-Sensitive Networking — TSN): Defines time-aware scheduling of Ethernet frames for deterministic sub-millisecond latency in industrial and automotive LANs.
  • 3GPP Release 15+: Specifies 5G URLLC (Ultra-Reliable Low-Latency Communication) targeting 1 ms air-interface latency and 99.9999% reliability for industrial IoT.
  • ETSI MEC (Multi-Access Edge Computing): Standardises edge compute deployment within mobile network infrastructure to enable latency-sensitive applications.
  • WebRTC W3C/IETF Standard: Joint specification for real-time peer-to-peer media in browsers, underpinning video conferencing, cloud gaming, and remote collaboration.
  • RDMA/InfiniBand (InfiniBand Trade Association): Kernel-bypass fabric standard for data-centre networks achieving sub-microsecond MPI and storage latencies.
  • Linux PREEMPT_RT: A long-standing patchset (mainlined in Linux 6.x) providing full kernel preemption and bounded interrupt latency for soft-real-time applications.
  • Time-Sensitive Networking (TSN): IEEE 802.1 family of standards for deterministic Ethernet in automotive (AUTOSAR) and industrial (OPC UA over TSN) contexts.

Historical Context

  • Low-latency engineering emerged from telecommunications and industrial control, where real-time response guarantees were embedded in standards for avionics, SCADA systems, and process control. The Internet’s packet-switched architecture introduced variable latency that required explicit management; MPLS traffic engineering and DiffServ emerged in the late 1990s to give operators control over latency-sensitive flows. Financial markets turbocharged the discipline through the high-frequency trading arms race of the 2000s, producing co-location services, kernel-bypass networking (DPDK), and FPGA-based matching engines. The 2010s brought mobile broadband and cloud gaming as mass-market drivers, and the 2020s introduced XR spatial computing and generative AI inference as new latency-sensitive frontiers demanding custom silicon and pervasive edge infrastructure.

Current Landscape (2026)

  • L4S (Low Latency, Low Loss, Scalable throughput; IETF RFC 9330/9331) moved from wired niche to mainstream mobile: it was folded into 3GPP Release 18 5G-Advanced, and in July 2025 T-Mobile became the first operator to unlock L4S across a wireless network at scale, exposing real-time radio-loading data to partners such as Apple and Vay.
  • A December 2025 SoftBank/Ericsson/Qualcomm field trial on SoftBank’s commercial 5G Standalone network in Tokyo combined L4S with Configured Uplink Grant and Rate-Controlled Scheduling to cut wireless-link latency by roughly 90% for smart-glasses XR streaming, with early ecosystem support also from Comcast, Apple (iOS 17) and MasOrange in Spain.
  • The Ultra Ethernet Consortium published Specification 1.0 on 11 June 2025 (refined to 1.0.1 in September 2025 and later 1.0.3), introducing the Ultra Ethernet Transport (UET) protocol with modern RDMA, packet spraying with NIC-side reordering, Link Layer Retry and new congestion control to deliver low tail latency for AI/HPC without requiring a lossless fabric.
  • AI-fabric silicon shipped through 2025: Broadcom’s 102.4 Tbps Tomahawk 6 (June 2025) offers a co-packaged-optics variant and Scale-Up Ethernet (SUE) positioned directly against NVLink, while NVIDIA’s Spectrum-X and Quantum-X Photonics CPO switches target ultra-low latency with Quantum-X InfiniBand availability early 2026 and Spectrum-X Ethernet in 2H 2026.
  • Scale-up interconnects consolidated around memory-semantic, sub-microsecond fabrics: NVIDIA announced NVLink Fusion (May 2025) to let third-party ASICs join its rack-scale domain, with NVLink 6 in the Rubin generation doubling GPU-to-GPU bandwidth to 3.6 TB/s, while the open UALink 1.0 specification defines a load/store fabric scaling to 1,024 accelerators.
  • The competitive latency picture as of 2026 sees InfiniBand NDR at roughly 1 microsecond (XDR switches targeting sub-500 nanoseconds) versus well-tuned Ethernet RoCE at about 1.5-2.5 microseconds, and by mid-2025 Ethernet had taken the lead in AI back-end deployments as UEC maturity and validated RoCE narrowed the gap.
  • Open challenges remain: UEC 1.0 hardware is still largely “UEC-ready” rather than fully compliant (for example AMD’s Pensando Pollara 400 NIC omits packet trimming and link-level CBFC), L4S needs end-to-end network-plus-endpoint support and is far from universally deployed, and co-packaged optics must prove field serviceability at 1.6 Tb/s densities before broad rollout.

References

Provenance