UniWAM Technical Report: Unified Mobile Manipulation via Mixed-Stream World-Action Modeling and Manipulation Anchor Pose Supervision

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of heterogeneous action spaces between navigation and interaction in mobile manipulation, which hinder unified policy learning, alongside the prohibitive cost of real-world data collection. To this end, this work proposes a Hybrid Flow World-Action Model featuring a shared backbone with decoupled encoding to support independent or parallel inference. Furthermore, it introduces a novel Manipulation Anchor Pose (MAP) supervision mechanism to resolve localization and orientation difficulties, integrated with an automated pipeline for scalable dataset construction. Experimental results demonstrate that the proposed method reduces positional error by 30.1% on the MAP-Bench benchmark and achieves state-of-the-art performance across 24 real-world tasks. The code and datasets have been made publicly available.
📝 Abstract
Mobile manipulation requires precise navigation to a manipulation-ready pose followed by reliable object interaction. These two stages differ in action spaces and visual requirements, which complicates unified policy learning. In addition, collecting diverse real-world navigation data with explicit manipulation-ready pose supervision remains costly and difficult to scale. We introduce UniWAM, a unified mixed-stream world-action model with separate action encoders and output heads for navigation and manipulation, sharing a common backbone. This design supports joint representation learning on independently sampled navigation and manipulation data. UniWAM supports independent inference for either stream and batch-parallel inference for both. We further introduce Manipulation Anchor Pose (MAP) supervision for where to stop and how to orient for manipulation. An automated pipeline constructs MAP-Data from large-scale 3D scenes, yielding over 1.5 million episodes and 7,500 hours. MAP-Data provides per-frame target-object bounding boxes and image-plane MAP coordinates as auxiliary navigation supervision. Together with projected end-effector trajectories for manipulation, these prediction targets provide stream-specific image-plane supervision for action learning from egocentric observations. With large-scale MAP-Data, UniWAM outperforms the strongest external baselines on our MAP-Bench by 30.1\% in position error and 44.0\% in heading error. Across 24 real-robot tasks, UniWAM achieves leading results in MAP navigation and mobile manipulation, with competitive manipulation performance. We have released code, data, and benchmark.
Problem

Research questions and friction points this paper is trying to address.

Mobile Manipulation
Unified Policy Learning
Navigation
Manipulation-ready Pose
Data Collection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unified Mobile Manipulation
Mixed-Stream World-Action Model
Manipulation Anchor Pose Supervision
Egocentric Observation
Large-scale 3D Data Generation
💼 Related Jobs
No related jobs found.
W
Wei Xue
Physical Superintelligence Lab, Fysics AI College of Intelligent Robotics and Advanced Manufacturing, Fudan University
K
Keliang Liu
Physical Superintelligence Lab, Fysics AI College of Intelligent Robotics and Advanced Manufacturing, Fudan University
M
Mingzhang Cui
Physical Superintelligence Lab, Fysics AI College of Intelligent Robotics and Advanced Manufacturing, Fudan University
J
Jinhua Xie
Physical Superintelligence Lab, Fysics AI College of Intelligent Robotics and Advanced Manufacturing, Fudan University
Jinjie Wei
Jinjie Wei
Fudan University
Large Language Model
J
Jianan Hou
Physical Superintelligence Lab, Fysics AI College of Intelligent Robotics and Advanced Manufacturing, Fudan University
J
Jingcheng Lu
Physical Superintelligence Lab, Fysics AI College of Intelligent Robotics and Advanced Manufacturing, Fudan University
Lintao Wang
Lintao Wang
The University of Sydney
character animationhuman motion understanding and generationlarge language modelai4science
K
Kaixiang Qiu
Physical Superintelligence Lab, Fysics AI College of Intelligent Robotics and Advanced Manufacturing, Fudan University
Yizhou Liu
Yizhou Liu
MIT
Dynamical systemsStatistical physicsPhysics of living systemsPhysics of AI
X
Xinghai Ye
Physical Superintelligence Lab, Fysics AI College of Intelligent Robotics and Advanced Manufacturing, Fudan University
J
Jinghang Han
Physical Superintelligence Lab, Fysics AI College of Intelligent Robotics and Advanced Manufacturing, Fudan University
Mingcheng Li
Mingcheng Li
Fudan University
J
Jie Gu
Physical Superintelligence Lab, Fysics AI College of Intelligent Robotics and Advanced Manufacturing, Fudan University
Shunli Wang
Shunli Wang
Fudan University
Computer visionAction quality assessment
Lihua Zhang
Lihua Zhang
Wuhan University
computational biologybioinformaticsdata mining
Dingkang Yang
Dingkang Yang
ByteDance
Multimodal LearningGenerative AIEmbodied AI