🤖 AI Summary
This study addresses the limitations of existing world action models, which lack explicit spatial understanding and suffer from architectural fragmentation due to multi-task supervision. This work proposes a unified framework that encodes RGB, action, and geometric information into a single video representation, enabling cross-modal temporal prediction via a unified diffusion Transformer. By introducing a deterministic video codec mechanism and eliminating task-specific head networks, the method achieves joint learning of structured perception and action through single-objective optimization. Experimental results demonstrate that the proposed approach attains a 52% closed-loop success rate on RLBench, doubling the baseline performance, while significantly outperforming cascaded schemes in future depth and segmentation prediction accuracy.
📝 Abstract
World action models (WAMs) have advanced robot control by predicting how observations and actions evolve over time. Despite this progress, RGB and action based future prediction does not explicitly address the spatial understanding needed for robot manipulation. Existing efforts often add a limited set of spatial prediction tasks through specialized heads or branches, leaving both the range of spatial supervision and the model architecture fragmented. We introduce XGenAct, a world action model that represents RGB observations, robot actions, metric depth, surface normals, and functional role segmentation as RGB videos through deterministic codecs. By sampling perception and action tasks during training, XGenAct uses one video diffusion transformer and one objective to learn temporal prediction across these spaces without modality specific learned heads. On held out RLBench tasks, structured perception training improves average closed loop success over RGB only training, and XGenAct achieves 52% success in the five task external comparison, versus 26% for the strongest evaluated baselines. It also predicts future depth and segmentation more accurately than the evaluated pipelines that generate RGB first and then apply a frozen perception expert.