OpenPose is an open-source, real-time multi-person pose estimation library developed at Carnegie Mellon University that simultaneously detects body, hand, face, and foot keypoints from RGB images and video using convolutional neural networks. It employs Part Affinity Fields (PAFs) — a set of 2D vector fields encoding the location and orientation of limb connections — enabling association of detected keypoints into individual skeletons without prior person detection. OpenPose established the part-affinity-field paradigm that underpins many subsequent human pose estimation systems.
Content
- OpenPose was developed at the Perceptual Computing Lab at Carnegie Mellon University, with the foundational paper “Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields” published at CVPR 2017 by Cao et al. Prior to OpenPose, multi-person pose estimation required first running a person detector and then estimating poses independently for each crop — a sequential, slow pipeline. OpenPose introduced a bottom-up approach: it detects all keypoints across all people simultaneously, then groups them using part affinity fields into individual skeletons, enabling real-time processing.
- Technically, OpenPose passes an image through a VGG-based or MobileNet backbone to produce two branches of outputs: confidence maps (heatmaps indicating keypoint locations) and part affinity fields (vector fields indicating limb directions). A bipartite graph matching algorithm then assembles the detected keypoints into complete skeleton instances. The system was extended to handle face landmarks (70 points), hand keypoints (21 per hand), and foot keypoints, providing a comprehensive 135-point body model. It runs at 22 fps on a single GPU for typical-resolution images with multiple people.
- The significance of OpenPose stems from democratising markerless human body tracking for research and application development. It enabled a generation of applications in action recognition, sports analytics, rehabilitation monitoring, dance and choreography analysis, physical therapy assessment, and human-robot collaboration. In AR/VR, OpenPose-derived pose data is used to drive avatar animation without dedicated motion capture suits. Its permissive open-source licence (academic, research use) facilitated widespread adoption and derivative work.
- By 2024-2025, OpenPose has been partially superseded by more efficient top-down systems (HRNet, ViTPose) and by transformer-based models that achieve higher accuracy on standard benchmarks. However, its bottom-up architecture remains valuable in dense crowd scenarios where the number of people is unknown. MediaPipe, Google’s production-grade successor framework, draws on similar principles but runs entirely on-device with hardware acceleration. Foundation model approaches are beginning to unify 2D pose, 3D pose, and mesh recovery in single models, pointing toward the next generation of human body understanding.