SuperNav: An Agentic Navigation System for Any Task in Any Scene

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited generalization of service robots in unfamiliar environments caused by fine-tuning multimodal large models, proposing a tuning-free universal navigation agent architecture. The approach leverages a pretrained multimodal large model exclusively for high-level decision-making, invoking low-level navigation skills through a unified visual point interface and a dedicated agent framework to execute actions. Precise control is achieved by integrating context management with a visual feedback loop. Experimental results demonstrate that the proposed system significantly outperforms four baseline methods on instance-level and multi-objective tasks. Furthermore, evaluations in both the HM3D simulation environment and on a real-world quadruped robot platform validate its strong robustness and broad adaptability across diverse scenarios and tasks.
📝 Abstract
General-purpose service robots need navigation systems that can handle diverse human requests in unfamiliar environments, combining task generality with scene generality. Some existing methods fine-tune multimodal large language models (MLLMs) to predict navigation actions, making their behavior dependent on the coverage of navigation training data and potentially limiting generalization to new requests and environments. Our key insight is to let the MLLM focus on interpreting requests, understanding scenes, and making decisions while preserving its general-purpose capabilities and delegating motion execution to navigation tools. To realize this idea, we introduce SuperNav, which equips a pretrained MLLM with a specialized agent harness without navigation-specific fine-tuning of the MLLM. Our harness supports these decisions with Navigation Skills, agent-oriented Tools for physical interaction, and task-progress and context management. A unified visual-point interface connects decision-making to motion by allowing the model to specify destinations directly in images and revise its decisions from execution feedback. Together, these components support sustained navigation across different task requirements and environments. SuperNav outperforms four evaluated baselines on instance-level, multi-object, and demand-driven tasks. Category-level evaluation on HM3D and deployment on a real quadruped robot further demonstrate its applicability across environments. Project Page: https://zju3dv.github.io/SuperNav/
Problem

Research questions and friction points this paper is trying to address.

robot navigation
general-purpose service robots
multimodal large language models
task generality
scene generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic Navigation
Multimodal Large Language Models
Visual-Point Interface
Zero-shot Generalization
Service Robots
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jinkai Zhang
Zhejiang University
J
Jingyi Xu
Zhejiang University
Y
Yuanhong Yu
Zhejiang University
Jiarui Guo
Jiarui Guo
Peking University
R
Ruizhen Hu
Shenzhen University
H
Hujun Bao
Zhejiang University
Xiaowei Zhou
Xiaowei Zhou
Professor of Computer Science, Zhejiang University
Computer VisionComputer Graphics
Sida Peng
Sida Peng
Zhejiang University
Computer VisionComputer Graphics