Dream4ACT: A Shared Visual Action Interface for Multi-Embodiment Video-Action Modeling

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges that cross-embodiment action vectors lack image-space structure, making it difficult to leverage spatiotemporal priors from video generation models under incomplete execution configurations. To this end, we propose a multi-embodiment video-action world model. The core innovation lies in designing a shared visual interface termed "action view," which employs URDF-based forward kinematics for unified representation, alongside a training-free multi-view recovery mechanism that eliminates the need for learning embodiment-specific decoders. By integrating a video autoencoder, a diffusion transformer, and masked flow matching, the framework enables forward dynamics, inverse dynamics, and joint generation. Extensive experiments demonstrate that our approach achieves an average success rate of 88.98% on RoboTwin 2.0 and an overall score of 65.66 on TriWorldBench, effectively supporting closed-loop manipulation and predictive control.
📝 Abstract
Video generation models (VGMs) offer strong spatiotemporal priors for embodied observation--action modeling. However, joint-space action vectors lack explicit image-space structure and vary in dimensionality and semantics across embodiments, making it challenging to directly leverage the rich spatiotemporal priors of VGMs. End-effector visualizations provide an alternative but do not specify the full articulated configuration needed for robot execution. We present Dream4ACT, a world model built for joint video-action modeling across embodiments. To unify action representations across embodiments, we introduce a shared visual action interface, called action views, which render target joint configurations from four prescribed virtual cameras using URDF-based forward kinematics. This shared visual representation preserves embodiment-specific articulated geometry while allowing observation and action sequences to share a video autoencoder and diffusion transformer. Through masked flow-matching, our model supports forward dynamics, inverse dynamics, and joint observation--action generation within a single jointly trained model by varying which future sequences are corrupted. To recover executable action sequences from predicted action views, we propose a training-free, URDF-constrained multiview recovery mechanism, without a learned embodiment-specific decoder. Dream4ACT achieves an average success rate of 88.98\% on RoboTwin~2.0 and an overall score of 65.66 on TriWorldBench, supporting effective closed-loop manipulation and competitive action-conditioned multiview prediction through the visual action interface.
Problem

Research questions and friction points this paper is trying to address.

video-action modeling
multi-embodiment
action representation
video generation models
joint-space actions
Innovation

Methods, ideas, or system contributions that make the work stand out.

shared visual action interface
multi-embodiment video-action modeling
masked flow-matching
URDF-constrained multiview recovery
diffusion transformer
🔎 Similar Papers
No similar papers found.
X
Xiangyu Zhu
Beijing Humanoid Robot Innovation Center, Beijing Institute of Technology
J
Jin Xu
Beijing Humanoid Robot Innovation Center
Y
Yue Guo
Beijing Humanoid Robot Innovation Center, Harbin Institute of Technology, Shenzhen
Xin Wu
Xin Wu
Beijing University of Posts and Telecommunications
Deep LearningObject DetectionRemote SensingFractional Fourier Transform
Y
Yifan Sun
Beijing Humanoid Robot Innovation Center, China University of Mining & Technology, Beijing
X
Xiancong Ren
Beijing Humanoid Robot Innovation Center
J
Jianxin Sun
Beijing Humanoid Robot Innovation Center
Y
Yong Dai
Beijing Humanoid Robot Innovation Center
X
Xiaozhu Ju
Beijing Humanoid Robot Innovation Center