Optical flow is the pattern of apparent motion of objects, surfaces, and edges in a visual scene between consecutive frames of video, caused by relative movement between the observer and the scene. It is computed as a dense or sparse 2D velocity field over the image plane and is used in computer vision for motion estimation, video interpolation, action recognition, and autonomous navigation. Classical algorithms (Horn-Schunck, Lucas-Kanade) and deep learning approaches (RAFT, FlowNet) constitute the main methodological lineages.

Content

  • The theoretical foundation of optical flow was established by Horn and Schunck (1981) and Lucas and Kanade (1981), who independently derived variational and feature-tracking formulations based on the brightness constancy constraint — the assumption that pixel intensity is preserved between frames. These classical methods remained the standard for two decades, supplemented by pyramid-based coarse-to-fine schemes (Anandan, 1989) and later by energy minimisation approaches that incorporated smoothness regularisation to handle occlusions and discontinuities.
  • Deep learning transformed optical flow estimation beginning with FlowNet (2015), the first end-to-end trained convolutional network for dense flow prediction. Subsequent architectures — SpyNet, PWC-Net, and RAFT (Recurrent All-Pairs Field Transforms, 2020) — achieved state-of-the-art performance by learning iterative refinement of dense correlation volumes. RAFT in particular established a new paradigm of computing all-pairs feature similarity and refining flow through recurrent update operators, achieving sub-pixel accuracy on benchmark datasets such as Sintel and KITTI.
  • Applications span autonomous vehicles (segmenting static background from moving objects), medical imaging (tracking cardiac wall motion), film post-production (motion vector extraction for re-timing and compositing), and video compression (where flow informs inter-frame prediction). In robotics, visual odometry systems use sparse optical flow to estimate camera ego-motion without external sensors. Video game engines have incorporated optical flow-based frame interpolation (DLSS Frame Generation, AMD FSR 3) to generate synthetic intermediate frames and increase perceived frame rates.
  • Between 2023 and 2025, optical flow has become a standard intermediate representation in generative video models such as Stable Video Diffusion and Sora, where consistent motion across frames is enforced using flow supervision or flow-warped attention. Event cameras — which detect per-pixel brightness changes asynchronously at microsecond resolution — have spurred a new class of event-based optical flow algorithms capable of handling extremely high-speed motion that frame-based cameras cannot capture. Integration of flow estimation into neural radiance field and 3D Gaussian Splatting pipelines has enabled dynamic scene reconstruction from monocular video.