CDP: Towards Robust Autoregressive Visuomotor Policy Learning via Causal Diffusion

📅 2025-06-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address weak policy generalization and frequent failures in long-horizon tasks on real robots—caused by poor observation quality and strict real-time constraints—this paper proposes Causal Diffusion Policy (CDP). CDP innovatively conditions the diffusion model on historical action sequences, enabling temporally consistent and context-aware action prediction. It introduces a KV caching mechanism to reuse attention key-value pairs, balancing inference robustness with low latency. Furthermore, CDP fuses vision-action multimodal signals within a Transformer-based causal diffusion architecture. Evaluated on both simulation and real-robot platforms, CDP significantly outperforms existing methods: it maintains high-precision 2D/3D dexterous manipulation even under degraded image quality, and demonstrates superior stability and generalization in target localization, grasp planning, and long-duration task execution.

Technology Category

Computer Vision: Diffusion Models for VisionHumans and AI: Human-Aware Planning and Behavior PredictionMachine Learning: Causal Learning

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingResponsible Web: Human-perceived consequences of algorithmic deployment on the web
📝 Abstract
Diffusion Policy (DP) enables robots to learn complex behaviors by imitating expert demonstrations through action diffusion. However, in practical applications, hardware limitations often degrade data quality, while real-time constraints restrict model inference to instantaneous state and scene observations. These limitations seriously reduce the efficacy of learning from expert demonstrations, resulting in failures in object localization, grasp planning, and long-horizon task execution. To address these challenges, we propose Causal Diffusion Policy (CDP), a novel transformer-based diffusion model that enhances action prediction by conditioning on historical action sequences, thereby enabling more coherent and context-aware visuomotor policy learning. To further mitigate the computational cost associated with autoregressive inference, a caching mechanism is also introduced to store attention key-value pairs from previous timesteps, substantially reducing redundant computations during execution. Extensive experiments in both simulated and real-world environments, spanning diverse 2D and 3D manipulation tasks, demonstrate that CDP uniquely leverages historical action sequences to achieve significantly higher accuracy than existing methods. Moreover, even when faced with degraded input observation quality, CDP maintains remarkable precision by reasoning through temporal continuity, which highlights its practical robustness for robotic control under realistic, imperfect conditions.
Problem

Research questions and friction points this paper is trying to address.

Improving robustness in visuomotor policy learning with degraded data
Reducing computational cost in autoregressive inference for real-time constraints
Enhancing action prediction accuracy using historical action sequences
Innovation

Methods, ideas, or system contributions that make the work stand out.

Transformer-based diffusion model for action prediction
Caching mechanism reduces autoregressive computation cost
Leverages historical actions for robust policy learning
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
J
Jiahua Ma
Sun Yat-sen University
Y
Yiran Qin
CUHK(SZ)
Y
Yixiong Li
Sun Yat-sen University
X
Xuanqi Liao
Sun Yat-sen University
Yulan Guo
Yulan Guo
Professor, Sun Yat-sen University
3D VisionMachine LearningRobotics
R
Ruimao Zhang
Sun Yat-sen University