RMMBench: A Comprehensive Benchmark for Robotic Mobile Manipulation

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of comprehensive evaluation benchmarks and limited metrics for assessing the embodied capabilities of vision-language models (VLMs) in robotic navigation and manipulation. To bridge this gap, we propose a unified evaluation framework that seamlessly integrates high-level planning with low-level control tasks for the first time. Specifically, we construct a "navigation-manipulation" benchmark suite comprising 70 scenes, enabling systematic assessment of fine-grained, long-horizon continuous spatial capabilities in mobile manipulation. Our empirical analysis reveals significant deficiencies in mainstream VLMs regarding complex spatial localization, underscoring the necessity of enhancing their spatial perception during extended interactions. Ultimately, this work provides critical infrastructure and identifies new research directions for advancing embodied artificial intelligence.
📝 Abstract
Although the advancement of vision-language models (VLMs) has endowed robots with enhanced environmental understanding and task reasoning, a comprehensive evaluation methodology is important to advance the integration of VLMs in robotic navigation and manipulation. However, current benchmarks lack a comprehensive method to evaluate diverse robotic tasks, and evaluation metrics remain relatively constrained, making it difficult to assess the embodied capabilities of VLMs in a thorough and fine-grained manner. To address this issue, we propose RMMBench, an evaluation benchmark that requires robots to understand language instructions and perform long-horizon tasks in continuous spaces. RMMBench seamlessly integrates high- and low-level embodied tasks into a unified framework, constructing a"navigation-manipulation"task suite comprising 70 canonical task scenarios that range from localized manipulation to long-horizon composite navigation. The results reveal that leading VLMs still face major challenges in spatial localization when performing mobile manipulation tasks, and also highlight the necessity of enhancing the spatial perception capability of robots during long-horizon interactions. RMMBench can be accessed at https://mxxq-stack.github.io/rmmbench-project/
Problem

Research questions and friction points this paper is trying to address.

Robotic Mobile Manipulation
Vision-Language Models
Benchmark
Embodied AI
Evaluation Metrics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Robotic Mobile Manipulation
Vision-Language Models
Benchmark
Embodied AI
Long-horizon Tasks
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Huapeng Li
School of Control Science and Engineering, Shandong University
F
Fuxiang Feng
School of Control Science and Engineering, Shandong University
J
Jinqiu Fan
School of Control Science and Engineering, Shandong University
S
Shuo Yang
School of Control Science and Engineering, Shandong University
F
Fengjiao Chen
Meituan Group
Xuezhi Cao
Xuezhi Cao
Meituan
Data MiningKnowledge GraphLLMs
R
Ran Song
School of Control Science and Engineering, Shandong University
W
Wei Zhang
School of Control Science and Engineering, Shandong University