JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation

๐Ÿ“… 2026-09-21
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
ๆœฌๆ–‡ๆๅ‡บJAMBๆจกๅž‹๏ผŒ้€š่ฟ‡่”ๅˆๅŽปๅ™ชๅŒ่‡‚ๅŠจไฝœๅ’Œๆœชๆฅ3D็‚น่ฝจ่ฟนๆฅ่งฃๅ†ณๅŒ่‡‚ๅ่ฐƒๆ“ไฝœไธญ็š„ๅ‡ ไฝ•ๅŽๆžœ้ข„ๆต‹้—ฎ้ข˜๏ผŒๆ˜พ่‘—ๆ้ซ˜ไบ†ไปปๅŠกๆˆๅŠŸ็އๅ’Œๆณ›ๅŒ–่ƒฝๅŠ›ใ€‚
๐Ÿ“ Abstract
Coordinated bimanual manipulation is challenging because the motion of either arm can alter the shared 3D scene and thereby affect the other arm. Yet most diffusion policies generate actions without explicitly modeling these future geometric consequences, while predictive variants typically use future state only as auxiliary supervision or fixed conditioning. We address this limitation by proposing JAMB, a diffusion policy that jointly denoises bimanual actions and future 3D point tracks. By allowing action and track hypotheses to evolve together within a shared Transformer, each can inform and refine the other throughout denoising. We further ground multimodal representations in a shared spatiotemporal coordinate system to facilitate geometry-aware interaction during joint denoising. We evaluate JAMB on diverse bimanual manipulation tasks in RoboTwin 2.0 and on a real-world robot, comparing it with action-only policies and alternative future-prediction approaches spanning different state representations and learning objectives. Across 16 simulation tasks, JAMB achieves an average success rate of 83.4%, outperforming the strongest baseline by 23.9 percentage points. On three real-world tasks, it outperforms the action-only and auxiliary geometry prediction methods by 50.0 and 21.2 percentage points, respectively. Beyond these performance gains, JAMB shows stronger generalization to cluttered scenes and out-of-distribution backgrounds than the evaluated baselines. Together, these results demonstrate the effectiveness of our joint action-motion modeling framework for coordinated bimanual manipulation. Our project website is available at https://jam-bimanual.github.io/
Problem

Research questions and friction points this paper is trying to address.

bimanual manipulation
diffusion policy
future geometric consequences
Innovation

Methods, ideas, or system contributions that make the work stand out.

Joint Action-Motion Diffusion
Bimanual Manipulation
Shared Transformer
Geometry-aware Interaction
Multimodal Representations
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
C
Chuyang Xiao
Robotics Institute, Carnegie Mellon University, Pittsburgh, PA 15213, USA
P
Peilin Meng
University of Michigan, Ann Arbor, MI 48109, USA
David Held
David Held
Associate Professor in the Robotics Institute, Carnegie Mellon University
RoboticsComputer VisionMachine LearningDeep LearningReinforcement Learning