Multi-View Stereo (MVS) is a computer vision technique that reconstructs dense 3D geometry from a set of overlapping 2D images captured from multiple camera positions. It extends traditional stereo matching by leveraging consistency across many viewpoints to estimate depth and surface detail at high resolution. MVS is a foundational component of photogrammetry pipelines, producing point clouds and textured meshes from photograph collections.

Content

  • MVS emerged as a formal research area in the early 2000s, motivated by the need to densify the sparse results of structure-from-motion pipelines. Seminal work by Seitz et al. (2006) on the Middlebury benchmark established standardised evaluation metrics and catalysed a decade of algorithmic competition. Approaches ranged from voxel carving and patch-based methods (PMVS, 2010) to depth-map fusion strategies, steadily improving both density and geometric accuracy.
  • The typical MVS pipeline proceeds in three stages: feature matching and camera pose estimation (inherited from SfM), per-view depth-map estimation via photometric consistency across neighbouring views, and depth-map fusion to produce a consistent dense point cloud or mesh. Plane-sweep stereo and PatchMatch-based algorithms dominate practical implementations. Key challenges include textureless surfaces, specular reflections, and occlusion boundaries, each of which breaks the assumption of consistent appearance across views.
  • In production, MVS is embedded within platforms such as RealityCapture, Agisoft Metashape, and open-source tools like OpenMVS and COLMAP. It underpins cultural heritage digitisation, construction site monitoring, film visual-effects pipelines, and autonomous vehicle map generation. Integration with Lidar sensors has become standard: LiDAR provides sparse but metrically accurate geometry that guides and corrects MVS depth estimates.
  • Between 2023 and 2025, neural approaches — particularly those inspired by NeRF and 3D Gaussian Splatting — have begun to supplant classical MVS for novel-view synthesis tasks, offering superior handling of view-dependent effects. However, classical MVS retains advantages in metric accuracy and interpretability for engineering and geospatial applications. Hybrid pipelines combining learned depth priors with classical geometric consistency checks represent the current frontier, narrowing the gap between photorealistic synthesis and metrically reliable reconstruction.