Spatially Conditioned Diffusion Policy: Learning Precise and Robust Manipulation with a Single RGB Camera

📅 2026-06-12
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of achieving high-precision and robust robotic manipulation using only a single global RGB camera, without relying on multi-view or wrist-mounted cameras. To this end, the authors propose a novel diffusion-based visuomotor policy that, for the first time, incorporates end-effector trajectories as spatial attention anchors within a diffusion framework. By integrating multi-scale visual encoding with trajectory-guided, point-level feature sampling, the model dynamically focuses on task-relevant regions. In simulation, the method significantly outperforms existing single-view approaches and matches the performance of multi-camera systems. Real-world experiments further demonstrate its robustness to visual distractions and its high manipulation accuracy.
📝 Abstract
Recent visual imitation learning systems have widely adopted multi-camera setups with wrist-mounted cameras as the de facto standard. However, manipulation from a single global view remains challenging, as the policy should capture fine-grained interaction details and identify task-relevant regions without local wrist views. To address this challenge, we present Spatially Conditioned Diffusion Policy (SCDP), a diffusion-based visuomotor policy that achieves precise and robust manipulation in a single-camera setting. Our key idea is that end-effector trajectories can serve as visual attention anchors that reflect task-relevant regions. Building on this idea, SCDP consists of two key components: (i) a visual encoder that produces multi-scale feature maps to capture both broader context and fine-grained visual features, and (ii) a spatial conditioning module that samples point-wise features along intermediate end-effector trajectories in the diffusion loop. Extensive simulation experiments show that SCDP consistently outperforms strong single-view baselines and achieves performance comparable to multi-camera baselines. Real-world experiments further demonstrate precise manipulation and robustness to visual distractors, highlighting the potential of single-camera imitation learning.
Problem

Research questions and friction points this paper is trying to address.

single-view manipulation
visual imitation learning
fine-grained interaction
task-relevant regions
RGB camera
Innovation

Methods, ideas, or system contributions that make the work stand out.

diffusion policy
single-view manipulation
spatial conditioning
visual imitation learning
end-effector trajectory
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Seoyoon Kim
Korea Advanced Institute of Science and Technology (KAIST)
Kanghyun Kim
Kanghyun Kim
Duke University
Computational ImagingMachine Learning
D
Dongwoo Ko
Neuromeka
Y
Yeong Jin Heo
Neuromeka
M
Min Jun Kim
Korea Advanced Institute of Science and Technology (KAIST)