SkeleWAM: Skeleton World-Action Modeling for Efficient Robotic Manipulation

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing world action models rely on video prediction, resulting in implicit and redundant geometric representations. This work proposes a world action model based on sparse 3D skeletons that replaces visual latent variables with online-constructed sparse 3D skeletons to provide explicit geometric supervision. By unifying geometric state representation without requiring visual reconstruction, the approach enables efficient action generation and future prediction. The method integrates RGB-D observations, proprioception, and a Medoid action consensus strategy for end-to-end training and inference. Evaluated on the LIBERO-Plus benchmark, the proposed model achieves an 85.9% success rate with only 57.1M parameters, outperforming Cosmos-Policy by 3.7 percentage points.
📝 Abstract
World action models (WAMs) combine robot action generation with future state prediction. Existing WAMs typically predict videos or learned visual latents, which represent interaction geometry only implicitly and may retain appearance information unrelated to control. We introduce SkeleWAM, a compact WAM that represents a manipulation scene as a sparse 3D skeleton composed of robot joints, object centers, and interaction points. Constructed online from current RGB-D observations and robot proprioception, the skeleton provides a unified geometric state for action generation and future skeleton prediction. Future skeleton prediction provides additional geometric supervision for action learning without requiring visual reconstruction. At inference, SkeleWAM generates actions directly from the current skeleton and language instruction, while Medoid Action Consensus (MAC) serves as an auxiliary consensus strategy for stochastic action samples. On LIBERO-Plus, SkeleWAM achieves an overall success rate of 85.9% with 57.1M parameters, outperforming Cosmos-Policy by 3.7 percentage points. These results demonstrate that sparse 3D robot--object structure provides an effective state space for robust and parameter-efficient world action learning. The project is available at https://skelewam-project.github.io/.
Problem

Research questions and friction points this paper is trying to address.

World Action Models
Robotic Manipulation
Interaction Geometry
Parameter Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Action Model
Sparse 3D Skeleton
Robotic Manipulation
Medoid Action Consensus
Parameter-efficient
💼 Related Jobs
No related jobs found.