A coordinate transformation is a mathematical mapping that converts the representation of a point, vector, or geometric object from one coordinate system or reference frame to another, preserving geometric relationships while expressing them in a new basis. In robotics and computer graphics, transformations are represented as homogeneous matrices, quaternions, or dual quaternions encoding rotation, translation, and scaling operations.
Content
- The mathematical theory of coordinate transformations descends from Euler’s work on rotation (1776), where he proved that any rigid-body rotation about a fixed point can be described as a single rotation about an axis (Euler’s theorem). Cauchy, Hamilton (quaternion algebra, 1843), and Cayley (matrix theory) provided the algebraic foundations. In robotics, Denavit and Hartenberg (1955) proposed a systematic 4x4 homogeneous transformation matrix convention for describing kinematic chains of rigid links, which became the standard in serial manipulator analysis. In computer graphics, affine transformation matrices were codified into the OpenGL pipeline during the 1990s.
- A 3D rigid-body transformation is represented by a 4x4 homogeneous matrix combining a 3x3 rotation matrix R and a 3x1 translation vector t in a single compact form, enabling composition of multiple transformations by matrix multiplication. Rotation matrices must satisfy R^T R = I and det(R) = +1 (they form the SO(3) Lie group). Euler angles (roll-pitch-yaw or ZYX convention) are intuitive but suffer from gimbal lock at singularities. Unit quaternions (elements of S^3) avoid gimbal lock and provide efficient interpolation (SLERP) for animation. Dual quaternions unify rotation and translation into a single algebraic object and are increasingly used in robotics for screw motion representation. The tf (transform) library in ROS manages a tree of time-stamped coordinate frames for heterogeneous sensor integration.
- Coordinate transformations matter because autonomous systems perceive their environment through sensors mounted at various locations on the robot body, each reporting data in its own local frame. Fusing lidar, camera, IMU, and GPS data requires expressing all measurements in a consistent global or body frame. In computer vision, camera projection matrices encode the perspective transformation from 3D world coordinates to 2D pixel coordinates. In augmented reality, real-time tracking maintains the transformation between device, world, and virtual content frames to achieve pixel-accurate overlay. In simulation, physics engines apply transformation trees to propagate forces and constraints through articulated body chains.
- In 2024-2025, differentiable coordinate transformations are central to neural 3D representations such as NeRF and Gaussian Splatting, where the rendering pipeline requires differentiable camera pose transformations for gradient-based optimisation of scene parameters. Learned pose estimation networks (PoseCNN, FoundPose) directly regress transformation parameters from images. Equivariant neural networks exploit coordinate transformation symmetries to build models whose outputs transform predictably with input pose, improving sample efficiency in robotics learning. SE(3)-equivariant networks for molecular property prediction in computational chemistry rely on the same mathematical foundations as robotics coordinate transforms.