🤖 AI Summary
This work investigates whether existing action-conditioned world models can generalize to unseen robot morphologies beyond mere visual memorization. To this end, the authors introduce XEWorld, a cross-embodiment evaluation benchmark that establishes, for the first time, an isolated-embodiment assessment paradigm. This framework systematically evaluates zero-shot and few-shot visual rendering capabilities of models when confronted with novel robots that share physical consistency but differ in embodiment structure. Experiments reveal that current models struggle to map abstract joint actions into coherent visual trajectories, relying heavily on visual similarity rather than kinematic or dynamic consistency for generalization. Furthermore, few-shot adaptation often leads to catastrophic forgetting of previously seen embodiments. These findings underscore the critical need for architectural innovations that explicitly disentangle appearance from physical dynamics.
📝 Abstract
Action-conditioned world models are promising learned simulators for robotic manipulation, yet evaluating them exclusively on training robots fails to reveal whether they capture physical dynamics or merely memorize visual patterns. To answer whether a model can faithfully render a robot it has never seen, we introduce XEWorld, a controlled cross-embodiment testbed for world models that isolates embodiments by evaluating held-out robots within physically identical scenes. Our systematic analysis uncovers a shared architectural bottleneck: current models act primarily as 2D visual pattern matchers whose generalization is governed by visual similarity rather than physical kinematic similarity. Driven by this limitation, they struggle to translate abstract numeric joint actions into coherent visual trajectories, and fail to predict dynamic visual changes from static initial observations. Consequently, successfully rendering an unseen embodiment zero-shot strictly requires heavily grounded cues, specifically pixel-space actions and explicit spatial-temporal alignment. Even when bypassing this zero-shot barrier via few-shot adaptation, the forced appearance recovery triggers catastrophic forgetting of seen embodiments. Together, these failures expose a critical inability to apply learned physical dynamics to novel visual appearances, highlighting that achieving true cross-embodiment generalization requires architectural innovations that decouple visual appearance from underlying physical dynamics.