XEWorld: Can Action-Conditioned World Models Generalize to Unseen Robot Embodiments?

📅 2026-08-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work investigates whether existing action-conditioned world models can generalize to unseen robot morphologies beyond mere visual memorization. To this end, the authors introduce XEWorld, a cross-embodiment evaluation benchmark that establishes, for the first time, an isolated-embodiment assessment paradigm. This framework systematically evaluates zero-shot and few-shot visual rendering capabilities of models when confronted with novel robots that share physical consistency but differ in embodiment structure. Experiments reveal that current models struggle to map abstract joint actions into coherent visual trajectories, relying heavily on visual similarity rather than kinematic or dynamic consistency for generalization. Furthermore, few-shot adaptation often leads to catastrophic forgetting of previously seen embodiments. These findings underscore the critical need for architectural innovations that explicitly disentangle appearance from physical dynamics.
📝 Abstract
Action-conditioned world models are promising learned simulators for robotic manipulation, yet evaluating them exclusively on training robots fails to reveal whether they capture physical dynamics or merely memorize visual patterns. To answer whether a model can faithfully render a robot it has never seen, we introduce XEWorld, a controlled cross-embodiment testbed for world models that isolates embodiments by evaluating held-out robots within physically identical scenes. Our systematic analysis uncovers a shared architectural bottleneck: current models act primarily as 2D visual pattern matchers whose generalization is governed by visual similarity rather than physical kinematic similarity. Driven by this limitation, they struggle to translate abstract numeric joint actions into coherent visual trajectories, and fail to predict dynamic visual changes from static initial observations. Consequently, successfully rendering an unseen embodiment zero-shot strictly requires heavily grounded cues, specifically pixel-space actions and explicit spatial-temporal alignment. Even when bypassing this zero-shot barrier via few-shot adaptation, the forced appearance recovery triggers catastrophic forgetting of seen embodiments. Together, these failures expose a critical inability to apply learned physical dynamics to novel visual appearances, highlighting that achieving true cross-embodiment generalization requires architectural innovations that decouple visual appearance from underlying physical dynamics.
Problem

Research questions and friction points this paper is trying to address.

world models
cross-embodiment generalization
robotic manipulation
physical dynamics
visual appearance
Innovation

Methods, ideas, or system contributions that make the work stand out.

world models
cross-embodiment generalization
action-conditioned simulation
visual dynamics decoupling
robotic manipulation
Y
Yixiang Chen
New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences
J
Jiabing Yang
New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences
Y
Yuan Xu
New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences
Q
Qisen Ma
New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences
Keji He
Keji He
SDU << CASIA & NUS
Cross-modal LearningEmbodied AI
Peiyan Li
Peiyan Li
Ludwig-Maximilians-Universität München
data mininggraph mining
K
Kai Wang
New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences
Z
Ziheng He
New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences
X
Xiangnan Wu
New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences
J
Jing Liu
FiveAges
N
Nianfeng Liu
FiveAges
Yan Huang
Yan Huang
Institute of Automation, Chinese Academy of Sciences
computer visiondeep learningmultimodal learning
Liang Wang
Liang Wang
National Lab of Pattern Recognition
Computer VisionPattern RecognitionMachine Learning