Matrix multiplication is the binary operation that combines two matrices to produce a third, where each entry of the result is the dot product of a row of the first matrix with a column of the second. It is the fundamental computational primitive of linear algebra and the dominant operation in deep learning, where dense layers, attention and convolutions all reduce to large matrix or tensor products executed on parallel hardware such as GPUs and TPUs.
Overview
- Matrix multiplication composes linear transformations: multiplying matrices corresponds to applying one linear map after another. Computationally it is the workhorse of numerical computing because it is highly regular, parallelisable and arithmetic-intensive, making it ideal for the throughput-oriented design of GPUs and dedicated accelerators. In deep learning almost every layer is, at its core, a matrix or tensor multiplication: dense layers multiply activations by weights, attention multiplies queries by keys and by values, and convolutions can be reformulated as matrix products. Optimising these operations dominates both training and inference performance.
- Because so much of computing reduces to matrix multiplication, decades of effort have gone into making it fast: cache-aware blocking, vectorised BLAS libraries, and accelerators that hardwire dense multiply-accumulate. In deep learning the operation appears in dense layers, attention and reformulated convolutions, so improvements to matrix multiplication translate almost directly into faster training and inference for the whole field.
History and context
- Matrix multiplication has been studied since the nineteenth century, but its computational importance exploded with numerical linear algebra and, later, deep learning. Algorithmic advances such as Strassen’s method and hardware such as GPUs and tensor cores have repeatedly reshaped how it is executed at scale.
Mechanisms
- Dot-product definition: each output element is the inner product of a row and a column of the inputs.
- Composition of transformations: chained matrix products represent sequences of linear maps.
- Computational intensity: high arithmetic-to-memory ratio that suits parallel accelerators.
- Tiling and blocking: cache- and memory-aware decompositions used in high-performance kernels.
- Tensor generalisation: batched and higher-order products that underpin modern deep-learning frameworks.
- Mixed precision: lower-precision arithmetic on tensor cores to raise throughput while controlling error.
Applications
- Forward and backward passes of neural networks expressed as matrix products.
- Attention mechanisms in transformers, dominated by query-key-value matrix multiplications.
- Scientific computing, simulation and signal processing kernels.
- Recommendation and embedding systems performing large dense and sparse products.
Challenges and considerations
- Memory bandwidth: feeding data to fast arithmetic units is often the real bottleneck.
- Numerical precision: low-precision multiplication boosts speed but must control accumulation error.
- Sparsity: exploiting sparse structure efficiently on dense-optimised hardware is hard.
- Scheduling: tiling and parallel decomposition must match the target hardware.
Examples
- A dense layer computing activations as a weight matrix times an input batch.
- Attention scoring queries against keys via a large matrix product.
- Reformulating convolution as an im2col matrix multiplication for GPU efficiency.