Manipulation is the capability of robotic or autonomous systems to physically interact with, grasp, reposition, and transform objects in an environment through controlled mechanical action. It integrates perception, kinematics, dynamics, and planning to enable precise, dexterous, and adaptive contact-rich tasks. Robotic manipulation encompasses the full pipeline from object detection and pose estimation through grasp planning, motion execution, and force-regulated contact control. As a foundational capability in intelligent systems, it bridges physical embodiment with higher-level task reasoning and is central to industrial automation, service robotics, and human-robot collaboration.
Overview
- Robotic manipulation translates high-level task goals into sequences of physical actions that produce purposeful changes in the world. Unlike Locomotion, which moves the robot itself, manipulation moves objects extrinsic to the robot body.
- The field draws from classical mechanics, control theory, computational geometry, and — increasingly — Machine Learning and Reinforcement Learning to cope with real-world variability in object shape, mass, friction, and clutter.
- A complete manipulation system must solve the perception-to-action loop: sensing the scene via Computer Vision and Sensor Fusion, estimating object state via Object Pose Estimation, generating feasible motions via Motion Planning, and executing them while monitoring contact forces.
- Industrial manipulation (welding, painting, assembly) has been mature for decades. Unstructured manipulation — grasping novel objects, adapting to unexpected contacts — remains an active research frontier.
Key Components
Mechanical Hardware
- End Effector — the terminal device (gripper, hand, tool) that makes physical contact with objects; may be parallel-jaw, multi-fingered, suction-based, or task-specific.
- Robotic Arm — the articulated kinematic chain that positions and orients the end effector in workspace; characterised by its degrees of freedom (DOF), reach, and payload.
- Actuator — the motor, pneumatic, or hydraulic device that drives each joint; actuator bandwidth and back-driveability critically affect force sensitivity.
- Tactile Sensing — arrays of pressure or deformation sensors embedded in fingertips or palm; provide rich contact information unavailable from proprioception alone.
Kinematics and Dynamics
- Forward Kinematics — maps joint angles to end-effector pose; analytic for standard geometries, numerical otherwise.
- Inverse Kinematics — the inverse problem: given a desired end-effector pose, compute joint angles; often admits multiple solutions and singularities.
- Robot Dynamics — governs how forces and torques produce accelerations; essential for high-speed or precision tasks and for whole-body force control.
- Jacobian Matrix — the geometric mapping between joint velocities and end-effector velocities; central to differential kinematics and resolved-rate control.
Planning
- Grasp Planning — selects stable grasp configurations from contact points on object geometry; quality metrics include grasp wrench space and robustness to perturbation.
- Motion Planning — finds collision-free joint-space or task-space paths from start to goal, typically using sampling-based planners (RRT, PRM) or optimisation-based methods.
- Trajectory Generation — converts discrete waypoints into smooth, dynamically feasible time-parameterised trajectories.
- Task and Motion Planning — integrates geometric motion planning with symbolic task planning to sequence multi-step manipulation actions.
- Collision Detection — real-time geometric checks against environment models; safety-critical for unstructured scenes.
Control
- Force Control — regulates contact force rather than position; essential whenever the robot makes and breaks contact with compliant or fragile objects.
- Impedance Control — models the robot end-effector as a mass-spring-damper; provides compliant behaviour without explicit force sensing.
- Admittance Control — the dual formulation: maps measured force inputs to velocity outputs; preferred when the environment is stiff.
- Visual Servoing — closes the control loop using image-based feedback to reduce reliance on precise calibration.
Perception
- Object Pose Estimation — determines the 6-DOF position and orientation of target objects from RGB-D images or Point Cloud data.
- Instance Segmentation — identifies and delineates individual object instances in cluttered scenes.
- Depth Sensing — provides metric distance maps via stereo cameras, structured light, or time-of-flight sensors; foundational for 3-D scene understanding.
- Sensor Fusion — combines proprioception, vision, and contact data into coherent state estimates robust to individual sensor noise.
Grasp Taxonomy
- Power grasps — whole-hand enveloping contact; maximise stability and payload; typical in industrial grippers.
- Precision grasps — fingertip-only contact; enable dexterous in-hand reorientation; require multi-finger hands.
- Non-prehensile manipulation — pushing, pivoting, sliding without gripping; useful for objects too large or fragile to grasp.
- In-hand manipulation — regrasping and reorienting an object within the hand fingers without releasing it; requires Dexterous Manipulation capability.
- Bimanual manipulation — coordinated use of two robotic arms or hands to handle large, deformable, or assembly-level objects.
- Deformable object manipulation — cloth, rope, food items; requires special representations since rigid-body assumptions break down.
Learning-Based Approaches
- Reinforcement Learning applied to manipulation trains policies end-to-end in simulation (MuJoCo, Isaac Sim) then transfers to real hardware via domain randomisation.
- Imitation Learning (behaviour cloning, DAgger) bootstraps from human teleoperation demonstrations; drastically reduces sample complexity relative to RL from scratch.
- Diffusion Policy and related generative-model approaches model the action distribution directly, capturing multi-modal behaviour in complex manipulation tasks.
- Foundation Models (large vision-language models) are increasingly used as high-level task planners that decompose manipulation goals into sub-skills.
- Sim-to-real transfer remains a central challenge: simulation cannot perfectly model friction, deformable contacts, and sensor noise.
Applications
Industrial Automation
- Assembly — inserting pegs, fastening bolts, joining sub-components on automotive and electronics production lines; requires sub-millimetre precision.
- Welding and painting — path-following manipulation tasks where end-effector orientation and speed uniformity are paramount.
- Bin picking — unstructured grasp of randomly oriented parts from containers; the canonical hard manipulation problem in industry.
- Palletising and depalletising — high-throughput pick-and-place for logistics and warehousing.
Service and Collaborative Robotics
- Human-Robot Collaboration — cobots such as the UR series and Franka Emika Panda work alongside humans; require safe, force-limited manipulation.
- Kitchen and domestic tasks — opening jars, folding laundry, loading dishwashers; exemplify unstructured dexterous manipulation at the frontier of capability.
- Retail automation — shelf stocking, item picking for e-commerce fulfilment (Amazon Kiva/Sparrow, Ocado).
Medical and Surgical Robotics
- Surgical Robotics (e.g., da Vinci system) requires extreme precision, miniaturised end effectors, and tremor filtering.
- Rehabilitation exoskeletons use manipulation principles for assisted limb movement.
- Laboratory automation (pipetting, sample handling) demands high repeatability at sub-millilitre scale.
Space and Hazardous Environments
- Satellite servicing manipulators (Canadarm, JEMRMS on ISS) operate in microgravity where reaction forces must be carefully managed.
- Nuclear decommissioning robots handle radioactive material remotely via Teleoperation with haptic feedback.
- Subsea manipulation for pipeline inspection and repair in high-pressure environments.
Extended Reality and Digital Twins
- Manipulation planning and verification in Digital Twin environments lets engineers validate robot programs offline before deployment.
- Spatial Computing interfaces enable operators to programme manipulation tasks intuitively via hand-tracking and gesture in AR/VR environments.
Standards & Context
- ISO 10218-1/2 — safety requirements for industrial robots and robot systems; governs speed and force limits relevant to collaborative manipulation.
- ISO/TS 15066 — specifies power and force limiting thresholds for human-robot collaborative operation; directly constrains compliant manipulation controller design.
- ISO 9283 — defines manipulator performance metrics: positioning accuracy, repeatability, path accuracy, and velocity.
- ROS (Robot Operating System) — the de-facto open middleware stack for manipulation research; MoveIt! is the canonical manipulation planning framework within ROS.
- URDF / SDF — XML formats for describing robot kinematic and dynamic parameters; universally used for manipulation simulation in Gazebo, MuJoCo, Isaac Sim.
- OpenRAVE — early open planning environment for manipulation; largely superseded by MoveIt! but historically influential.
- Research benchmarks: YCB Object and Model Set (Yale-CMU-Berkeley), OCRTOC challenge, and the Real Robot Challenge provide standardised evaluation for manipulation systems.
Semantic Classification
Current Landscape (2026)
- Vision-language-action (VLA) flow models are now the dominant manipulation paradigm: Physical Intelligence’s pi-0 (October 2024) and its successor pi-0.5 (April 2025) demonstrated, for the first time, end-to-end long-horizon dexterous skills such as cleaning kitchens and bedrooms in entirely unseen homes, using co-training on heterogeneous web, multi-robot and semantic-subtask data plus “knowledge insulation”.
- Open-weight foundation models have democratised the field: Physical Intelligence’s openpi release (with PyTorch support added September 2025) ships pi-0, pi-0-FAST and pi-0.5 checkpoints pre-trained on 10k+ hours of robot data, alongside Stanford/Berkeley OpenVLA (7B) and Berkeley Octo, all standardising on the Hugging Face LeRobot data schema.
- NVIDIA’s Isaac GR00T line has set the pace for humanoid manipulation: GR00T N1 (March 2025), N1.5 (11 June 2025) and later N1.6 progressively broadened cross-embodiment coverage (Fourier GR-1, Unitree G1, 1X Neo, AgiBot Genie-1), paired with the Cosmos/GR00T-Dreams world-foundation-model pipeline that synthesises “neural trajectory” training data from a single image and prompt.
- Dexterous end-effectors took a step change: Tesla revealed Optimus Gen 3 hands on 17 February 2026 with 22 DoF per hand and 25 forearm-mounted tendon-driven actuators (50 total), targeting 3,000+ discrete tasks, while Figure 02 fields a 16-DoF hand with 3-gram fingertip force sensing and Sanctuary AI Phoenix uses a 20-21 DoF hydraulic hand.
- Commercial deployment is real but uneven: Agility Robotics’ Digit is shipping fleets to Amazon and GXO Logistics from its Salem RoboFab facility, whereas Tesla’s Optimus (a few hundred units as of Q2 2026 earnings, 22 July) remains in a training-data-collection phase with productive factory work still unverified by third parties.
- Synthetic data and world models have become central to closing the data gap, with sim-to-real pipelines (Isaac Sim/Isaac Lab, Cosmos Predict-2, inverse-dynamics action labelling) and NVIDIA’s open Physical AI dataset — the most-downloaded robotics dataset on Hugging Face — supplying the trajectories that pure real-world teleoperation cannot scale to.
- Open frontier challenges as of 2026 remain generalisation to genuinely unstructured environments, human-speed execution, reliable multi-step autonomy without supervision, robust tactile/force feedback, and safety verification — reflected in research directions such as SafeVLA and the persistent gap between demonstrated design intent and verified factory throughput.
References
-
- Physical Intelligence (2025). pi-0.5: a Vision-Language-Action Model with Open-World Generalization. https://www.pi.website/blog/pi05
-
- Black, K. et al. / Physical Intelligence (2024). pi-0: A Vision-Language-Action Flow Model for General Robot Control. https://arxiv.org/html/2410.24164v1
-
- NVIDIA Research (2025). GR00T N1.5: An Improved Open Foundation Model for Generalist Humanoid Robots. https://research.nvidia.com/labs/gear/gr00t-n1_5/
-
- NVIDIA (2025). Enhance Robot Learning with Synthetic Trajectory Data Generated by World Foundation Models (Isaac GR00T-Dreams). https://developer.nvidia.com/blog/enhance-robot-learning-with-synthetic-trajectory-data-generated-by-world-foundation-models/
-
- Wikipedia contributors (2026). Humanoid hand. https://en.wikipedia.org/wiki/Humanoid_hand
-
- Robozaps (2026). Tesla Optimus vs Agility Robotics Digit: Full 2026 Comparison. https://blog.robozaps.com/b/tesla-optimus-vs-agility-robotics-digit