🤖 AI Summary
This study addresses the challenges of heterogeneous action spaces between navigation and interaction in mobile manipulation, which hinder unified policy learning, alongside the prohibitive cost of real-world data collection. To this end, this work proposes a Hybrid Flow World-Action Model featuring a shared backbone with decoupled encoding to support independent or parallel inference. Furthermore, it introduces a novel Manipulation Anchor Pose (MAP) supervision mechanism to resolve localization and orientation difficulties, integrated with an automated pipeline for scalable dataset construction. Experimental results demonstrate that the proposed method reduces positional error by 30.1% on the MAP-Bench benchmark and achieves state-of-the-art performance across 24 real-world tasks. The code and datasets have been made publicly available.
📝 Abstract
Mobile manipulation requires precise navigation to a manipulation-ready pose followed by reliable object interaction. These two stages differ in action spaces and visual requirements, which complicates unified policy learning. In addition, collecting diverse real-world navigation data with explicit manipulation-ready pose supervision remains costly and difficult to scale. We introduce UniWAM, a unified mixed-stream world-action model with separate action encoders and output heads for navigation and manipulation, sharing a common backbone. This design supports joint representation learning on independently sampled navigation and manipulation data. UniWAM supports independent inference for either stream and batch-parallel inference for both. We further introduce Manipulation Anchor Pose (MAP) supervision for where to stop and how to orient for manipulation. An automated pipeline constructs MAP-Data from large-scale 3D scenes, yielding over 1.5 million episodes and 7,500 hours. MAP-Data provides per-frame target-object bounding boxes and image-plane MAP coordinates as auxiliary navigation supervision. Together with projected end-effector trajectories for manipulation, these prediction targets provide stream-specific image-plane supervision for action learning from egocentric observations. With large-scale MAP-Data, UniWAM outperforms the strongest external baselines on our MAP-Bench by 30.1\% in position error and 44.0\% in heading error. Across 24 real-robot tasks, UniWAM achieves leading results in MAP navigation and mobile manipulation, with competitive manipulation performance. We have released code, data, and benchmark.