KineWorld: Action-Induced Transport Fields for Embodied World Modeling

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the misalignment between visual generation objectives and interaction requirements in existing embodied world models, as well as their neglect of critical sparse spatial variations. To this end, we propose KineWorld, a framework that incorporates robot kinematics into the spatial allocation of generative supervision. Specifically, it introduces Kinematic Transport Lifting (KTL) to construct camera-aligned transport fields and proposes Transport-Aware World Diffusion (TAWD) for reweighted flow matching. Combined with video latent grid calibration and a hybrid distribution normalization mechanism, KineWorld enables precise prediction of action consequences. Evaluated on RoboTwin 2.0, the framework achieves a single-view EWMScore-P of 68.95 and a multi-view TWB-Score of 54.82, validating a paradigm shift from appearance fitting toward action consequence modeling.
📝 Abstract
Embodied world models predict the visual consequences of candidate actions before execution. However, existing action-conditioned world models often adopt uniformly weighted visual generation objectives that can be misaligned with embodied prediction needs. Even with explicit motion conditioning, these objectives can underemphasize spatially sparse changes that are critical to interaction. We propose KineWorld, a transport-aware world-modeling framework that extends robot kinematics from motion conditioning to the spatial allocation of generative supervision. Kinematic Transport Lifting (KTL) constructs renderer-derived, camera-aligned transport fields from commanded robot motion. Transport-Aware World Diffusion (TAWD) calibrates their motion support on the video-latent grid and reweights future-RGB flow matching through a normalized mixture of uniform and transport-focused distributions. We train KineWorld using ALOHA-AgileX bimanual manipulation data from RoboTwin 2.0. KineWorld achieves an EWMScore-P of 68.95 in single-view evaluation and a TWB-Score of 54.82 in multi-view evaluation. These results support a shift from appearance fitting toward action-consequence modeling for embodied decision-making.
Problem

Research questions and friction points this paper is trying to address.

embodied world models
action-conditioned prediction
visual generation objectives
spatially sparse changes
action-consequence modeling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Embodied World Modeling
Kinematic Transport Lifting
Transport-Aware Diffusion
Flow Matching
Action-Consequence Modeling
🔎 Similar Papers
No similar papers found.