Score
Design, build, and train diffusion-based stochastic control policies (DDPM-style) that generate temporally coherent, multimodal action sequences or pose–action chunks conditioned on observation windows, keyposes, chronoflow/keypoint signals, or other contextual information. Implement and evaluate conditioning and sampling procedures, co-train perception representations with action prediction, and analyze closed-loop behavior, robustness, and generalization of the learned visuomotor policies.
This work addresses three core challenges in robotic manipulation: difficulty in modeling multimodal action distributions, poor robustness to high-dimensional input-output spaces, and scarcity of real-world demonstration data. To this end, we systematically investigate diffusion models for grasp learning, trajectory planning, and data augmentation. We propose a vision-action joint modeling framework that integrates denoising diffusion probabilistic models (DDPMs) with score distillation sampling (SDS), augmented by conditional generation and simulation-to-real domain-cooperative enhancement. We introduce the first taxonomy and evaluation benchmark specifically designed for diffusion models in robotic manipulation. Empirical results demonstrate significant improvements over conventional imitation learning and reinforcement learning paradigms in high-dimensional robustness, few-shot generalization, and cross-modal alignment. The study further identifies scalability, real-time inference efficiency, and physical consistency as three critical directions for future advancement.
This work addresses the limited responsiveness and poor adaptability of diffusion-based policies in dynamic environments, which often lead to task failure. To overcome these limitations, the authors propose the DCDP framework, which integrates chunked action generation with a training-free, real-time closed-loop correction mechanism to significantly enhance robotic responsiveness and adaptability in dynamic scenarios. The approach innovatively combines a self-supervised dynamic feature encoder, cross-attention fusion, and an asymmetric action encoding-decoding architecture to enable efficient online adjustment. Evaluated on the dynamic PushT simulation benchmark, DCDP improves task adaptability by 19% with only a 5% increase in computational overhead and supports plug-and-play deployment on real-world robotic manipulation tasks.
This work addresses the adaptive challenges posed by non-stationary task dynamics and evolving goals in visual reinforcement learning. We introduce diffusion policies to this setting for the first time, proposing an iterative denoising-based diffusion policy framework that directly generates temporally coherent and context-aware action sequences from high-dimensional visual observations. Coupled with a lightweight visual encoder, the framework enables efficient online adaptation. Evaluated on non-stationary benchmarks—including Procgen and PointMaze—our method outperforms PPO and DQN, achieving +12.7%–28.3% gains in average and peak reward while reducing action variance by 36.5%, demonstrating robust adaptation to drifting dynamics and shifting task objectives. Our core contribution is the pioneering integration of diffusion models into non-stationary visual RL, unifying efficient exploration with stable online adaptation.
To address weak policy generalization and frequent failures in long-horizon tasks on real robots—caused by poor observation quality and strict real-time constraints—this paper proposes Causal Diffusion Policy (CDP). CDP innovatively conditions the diffusion model on historical action sequences, enabling temporally consistent and context-aware action prediction. It introduces a KV caching mechanism to reuse attention key-value pairs, balancing inference robustness with low latency. Furthermore, CDP fuses vision-action multimodal signals within a Transformer-based causal diffusion architecture. Evaluated on both simulation and real-robot platforms, CDP significantly outperforms existing methods: it maintains high-precision 2D/3D dexterous manipulation even under degraded image quality, and demonstrates superior stability and generalization in target localization, grasp planning, and long-duration task execution.
This work addresses the challenge of achieving high-precision and robust robotic manipulation using only a single global RGB camera, without relying on multi-view or wrist-mounted cameras. To this end, the authors propose a novel diffusion-based visuomotor policy that, for the first time, incorporates end-effector trajectories as spatial attention anchors within a diffusion framework. By integrating multi-scale visual encoding with trajectory-guided, point-level feature sampling, the model dynamically focuses on task-relevant regions. In simulation, the method significantly outperforms existing single-view approaches and matches the performance of multi-camera systems. Real-world experiments further demonstrate its robustness to visual distractions and its high manipulation accuracy.
Diffusion policies excel at modeling high-dimensional, multimodal behaviors but suffer from inherent stochasticity and offline training paradigms, limiting their applicability to real-time control under dynamic, previously unseen constraints. To address this, we propose a constraint-aware diffusion predictive control framework that, for the first time, integrates constraint tightening and model-driven dynamical projection directly into the diffusion denoising process—enabling online satisfaction of out-of-distribution constraints. Our method builds upon a pre-trained trajectory diffusion model and jointly incorporates system dynamics projection, robust constraint tightening, and iterative optimization. Evaluated in robotic arm simulations, the approach significantly improves constraint satisfaction rates for previously unseen state and action constraints while preserving task performance. It consistently outperforms both existing diffusion-based policies and conventional model predictive control (MPC) methods in terms of constraint adherence, tracking accuracy, and computational efficiency.
Traditional diffusion policies struggle to precisely control behavior generation and lack effective modeling of semantically similar trajectories. This work proposes Parametric Diffusion Policy (PDP), which constructs a low-dimensional continuous behavioral manifold to transform the diffusion process into a controllable generative mechanism guided by semantic parameters. This approach enables behavior interpolation and zero-shot adaptation without updating policy weights. By integrating diffusion policy learning, behavioral manifold embedding, and semantic-aware trajectory representation, PDP significantly outperforms existing diffusion-based methods in multimodal robotic tasks—both simulated and real—demonstrating superior control precision and adaptability, particularly in synthesizing novel behaviors.
This work addresses the high inference latency of diffusion policies in visuomotor robot control, which stems from iterative denoising and hinders the simultaneous achievement of real-time performance and high-quality actions. To overcome this limitation, the authors propose STEP, a method that leverages a lightweight spatiotemporal consistency predictor to generate high-quality warm-start actions and incorporates a velocity-aware perturbation injection mechanism to substantially improve both inference efficiency and execution stability without compromising generative capability. Theoretical analysis shows that the predictor induces a local contraction mapping, ensuring convergence of action errors. Evaluated across nine simulation benchmarks and two real-world tasks, STEP achieves an average success rate improvement of 21.6% over BRIDGER on RoboMimic and 27.5% over DDIM in physical experiments using only two denoising steps, significantly advancing the Pareto frontier between latency and success rate.
Existing diffusion models for robotic motion generation struggle to simultaneously model temporal dependencies and achieve real-time inference, often constrained by short-horizon synthesis or high latency from multi-step sampling. This work proposes distilling diffusion models into the parameter space of Probabilistic Dynamic Movement Primitives (ProDMPs), enabling single-step consistency distillation to generate full-horizon trajectories with realistic acceleration and deceleration dynamics at high speed. The method uniquely supports one-step generation of complete, temporally structured motion primitives, preserving dynamic characteristics while entirely eliminating the multi-step inference bottleneck. Experiments on MetaWorld and ManiSkill benchmarks demonstrate a 10× speedup over MPD and a 7× improvement over action chunking strategies, with comparable or higher task success rates, and enable real-time interception of fast-moving aerial objects.
This work addresses the challenge that diffusion-based policies often fail to discover rare yet effective behavioral patterns under scarce demonstration data, frequently converging to suboptimal solutions or generating infeasible trajectories. To overcome this limitation, the authors propose a novel framework integrating a Feynman–Kac corrector with a learnable guidance potential, which systematically steers the diffusion process toward under-explored feasible regions of the trajectory space. By coupling this guided exploration with sample-based trajectory optimization and an iterative retraining mechanism, the method continuously refines and reuses newly discovered behaviors. Empirical results demonstrate that the approach substantially enhances policy diversity, feasibility, and generalization, consistently uncovering novel and effective strategies across diverse manipulation tasks—surpassing the capabilities of conventional sampling and reinforcement learning methods.
本文探讨了通过控制扩散过程来解决机器人学习中的分布形状和轨迹问题,利用受控扩散和遍历控制方法优化机器人行为。