KineWorld: Action-Induced Transport Fields for Embodied World Modeling
This study addresses the misalignment between visual generation objectives and interaction requirements in existing embodied world models, as well as their neglect of critical sparse spatial variations. To this end, we propose KineWorld, a framework that incorporates robot kinematics into the spatial allocation of generative supervision. Specifically, it introduces Kinematic Transport Lifting (KTL) to construct camera-aligned transport fields and proposes Transport-Aware World Diffusion (TAWD) for reweighted flow matching. Combined with video latent grid calibration and a hybrid distribution normalization mechanism, KineWorld enables precise prediction of action consequences. Evaluated on RoboTwin 2.0, the framework achieves a single-view EWMScore-P of 68.95 and a multi-view TWB-Score of 54.82, validating a paradigm shift from appearance fitting toward action consequence modeling.