Camera Intrinsics are the internal optical and geometric parameters of a camera that define the mathematical relationship between 3D points in the camera’s coordinate frame and their 2D projections onto the image sensor. The intrinsic parameter matrix encodes focal length in pixel units along each image axis, the principal point (optical axis intersection with the sensor), and skew, whilst associated distortion coefficients correct for lens aberrations that cause deviations from the ideal pinhole projection model.

Content

  • The mathematical model of camera intrinsics derives from the pinhole camera model, a projective geometry idealisation in which a 3D scene point is projected through a single point (the optical centre) onto an image plane at distance (the focal length). This model was formalised for computer vision applications by Hartley and Zisserman in their canonical text Multiple View Geometry in Computer Vision (2000), building on earlier photogrammetric work in surveying and aerial mapping. The decomposition of the projection matrix P = K[R|t] into intrinsic matrix K and extrinsic rotation R and translation t provides the mathematical foundation for most multi-view reconstruction algorithms.
  • Technically, the intrinsic matrix where and (d_x, d_y being the physical pixel dimensions). For most modern cameras, skew is negligible. Distortion is modelled by the Brown-Conrady model with radial coefficients and tangential coefficients . Calibration solves for these parameters by minimising reprojection error over a set of images of a planar calibration target (checkerboard) from multiple viewpoints, using Zhang’s method. Modern implementations (OpenCV, MATLAB Computer Vision Toolbox) handle this via non-linear least-squares (Levenberg-Marquardt).
  • Camera intrinsics matter across autonomous vehicles (LiDAR-camera fusion requires precise intrinsic calibration of each camera in the array), extended reality headsets (display-camera calibration determines geometric correction for see-through AR), robotic manipulation (grasp point estimation from RGB-D cameras requires metric 3D reconstruction), medical imaging (endoscope calibration for surgical guidance), and satellite remote sensing (geometric correction of pushbroom sensor imagery). Each application domain has its own calibration field practices, accuracy requirements, and repeatability standards.
  • In 2024–2025, deep learning approaches to intrinsic calibration are supplementing classical methods. Neural networks trained on large image datasets can estimate approximate focal length and distortion from single images, enabling rapid intrinsic bootstrapping without a physical calibration target. Learning-based undistortion models bypass explicit polynomial parameterisation altogether, operating directly in pixel space. For fish-eye and catadioptric cameras — where the polynomial distortion model breaks down — generalised unified camera models and neural implicit representations are being adopted. Factory calibration with robotic precision fixtures is increasingly combined with online self-calibration to track thermal drift and focus changes during operation.