Diffusion Policy is a class of robot learning algorithms that represent robot action sequences as the output of a conditional denoising diffusion process, treating action prediction as iterative noise removal conditioned on sensor observations rather than as direct regression or classification. By leveraging the expressiveness of diffusion models to capture multi-modal action distributions, Diffusion Policy can represent one-to-many mappings from observation to action — a critical capability for dexterous manipulation tasks where multiple valid action trajectories exist. The approach, introduced by Chi et al. (2023), achieves state-of-the-art performance on imitation learning benchmarks and generalises across diverse manipulation settings.
Content
- Diffusion Policy was introduced by Cheng Chi, Siyuan Feng, Yilun Du, Zhenjia Xu, Eric Cousineau, Benjamin Burchfiel, and Shuran Song in a 2023 paper that demonstrated the limitations of regression-based behavioural cloning when action distributions are multi-modal. The core insight was that diffusion models, which had proven highly expressive for image and audio generation, could be repurposed as policy representations where the “noise” is in action space and the denoising network conditions on observation embeddings from cameras and proprioception sensors.
- The training procedure collects demonstrations — typically via kinesthetic teaching or teleoperation — as (observation, action) pairs. A denoising network is trained to predict the noise added to ground-truth action sequences at varying noise levels (the standard DDPM/DDIM objective). At inference, the policy begins from Gaussian noise in action space and iteratively denoises it over K steps (typically 10-100), conditioned on the current observation, yielding a full action chunk representing a short-horizon trajectory. This chunk-based action prediction, inspired by the temporal consistency of diffusion outputs, also helps mitigate compounding errors relative to single-step behavioural cloning.
- Diffusion Policy achieves strong empirical results on contact-rich manipulation benchmarks including push-T, robomimic, and real-world manipulation tasks involving deformable objects, granular materials, and multi-step assembly. It handles bimanual coordination and generalises across object poses and scene configurations with relatively modest data requirements (tens to hundreds of demonstrations). The CNN-based and transformer-based variants of the denoising network offer different trade-offs between inference speed and expressiveness. Consistency policies and flow matching reformulations have since been proposed to reduce the inference-time computational burden.
- By 2024-2025 diffusion-based robot policies are central to the foundational robot learning agenda pursued by groups at Stanford, CMU, MIT, Columbia, and major industrial labs including Google DeepMind and Physical Intelligence (Pi). They are being scaled to vision-language-conditioned generalist robot systems, integrated with large vision-language models for instruction following, and deployed on humanoid platforms. Flow Matching as an alternative training objective is gaining traction for its simpler training dynamics and faster inference, and hardware-in-the-loop fine-tuning methods are making diffusion policies practical for manufacturing and logistics deployment.