MobiAgent: Dual-Loop Recursive Policy Self-Improvement for Long-Horizon Mobile Manipulation

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of accumulated errors and capacity interference between mobile base and manipulator control in long-horizon mobile manipulation. We propose a dual-loop agent framework based on vision-language models (VLMs). The inner loop achieves robust execution by planning composable atomic skills via VLMs, while the outer loop employs an unsupervised lifelong learning mechanism to automatically segment and validate data, recursively refining the skill library and overcoming the rigidity of conventional hierarchical agents. This approach outperforms baselines by 22.5% on BEHAVIOR-1K and improves success rates on RoboCasa from 7.5% to 27.5%, significantly enhancing the system's fault recovery capabilities.
📝 Abstract
Long-horizon mobile manipulation presents significant challenges due to compounding execution errors and capacity interference between locomotion and arm control. While recent Vision-Language-Action models excel at short-horizon tasks, they lack the hierarchical reasoning required for multi-stage objectives. Furthermore, existing hierarchical agents suffer from rigid sub-task mapping, inflexible replanning, and a lack of continuous learning. To address these limitations, we introduce MobiAgent, a dual-loop agentic framework that bridges robust deployment execution and recursive policy self-improvement. During deployment, the Inner Loop decouples high-level reasoning from low-level control through highly composable atomic skills. It employs Vision-Language models for receding-horizon planning and visual reflection, dynamically composing skills to ensure robust error recovery. These skills are executed by specialized flow-matching experts that share a unified VLM backbone, maximizing reusability while mitigating capacity interference. Concurrently, the Outer Loop drives automated lifelong learning by autonomously segmenting and verifying deployment rollouts, clustering them to discover atomic skills, and continuously fine-tuning the skill library without human annotations. Evaluations on RoboCasa, BEHAVIOR-1K, and real-world tasks demonstrate the effectiveness of MobiAgent. It outperforms $π_{0.5}$-TA by 22.5 percentage points on BEHAVIOR-1K and enables robust recovery from execution failures. Through autonomous data recycling, success improves from 7.50% to 27.50% on RoboCasa and from 32.5% to 57.5% on Astribot S1.
Problem

Research questions and friction points this paper is trying to address.

Long-horizon mobile manipulation
Compounding execution errors
Capacity interference
Hierarchical reasoning
Continuous learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mobile Manipulation
Dual-Loop Framework
Vision-Language Models
Lifelong Learning
Flow Matching
🔎 Similar Papers