Memory bandwidth is the rate at which data can be read from or written to memory, typically measured in gigabytes per second. It is a primary performance constraint for data-intensive workloads such as deep-learning inference and real-time rendering, where compute units stall waiting for data. High-bandwidth memory technologies are deployed precisely to relieve this bottleneck on accelerators and edge hardware.

Content

  • Many modern workloads are memory-bound rather than compute-bound, so techniques like operator fusion, quantisation, and on-chip caching aim to reduce memory traffic. High-bandwidth memory (HBM) stacks and wide buses raise the ceiling, but bandwidth per FLOP continues to shape the efficiency of AI and graphics systems.