OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping

📅 2025-08-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the challenge of precisely grounding open-vocabulary, free-form natural language instructions to 3D scene instances in open-world embodied agents, this paper proposes a zero-shot open-vocabulary vision-language grounding method. Methodologically, it integrates (1) a structure-semantic consistency constraint to ensure geometric and semantic alignment of multi-view instance representations; (2) an LLM-assisted instruction-to-instance alignment module for fine-grained semantic understanding; and (3) a unified framework combining vision-language models, 3D instance aggregation, explicit geometric modeling, and LLM-driven spatial contextual reasoning. Experiments on ScanNet200 and Matterport3D demonstrate substantial improvements over state-of-the-art methods in semantic mapping and instruction-based target retrieval. Notably, our approach achieves, for the first time, open-vocabulary, instance-level instruction grounding without requiring category priors or labeled training data.

Technology Category

Natural Language Processing: Language Grounding & Multi-modal NLPComputer Vision: 3D Computer VisionMachine Learning: Large Multimodal Models (LMMs)

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Grounding natural language instructions to visual observations is fundamental for embodied agents operating in open-world environments. Recent advances in visual-language mapping have enabled generalizable semantic representations by leveraging vision-language models (VLMs). However, these methods often fall short in aligning free-form language commands with specific scene instances, due to limitations in both instance-level semantic consistency and instruction interpretation. We present OpenMap, a zero-shot open-vocabulary visual-language map designed for accurate instruction grounding in navigation tasks. To address semantic inconsistencies across views, we introduce a Structural-Semantic Consensus constraint that jointly considers global geometric structure and vision-language similarity to guide robust 3D instance-level aggregation. To improve instruction interpretation, we propose an LLM-assisted Instruction-to-Instance Grounding module that enables fine-grained instance selection by incorporating spatial context and expressive target descriptions. We evaluate OpenMap on ScanNet200 and Matterport3D, covering both semantic mapping and instruction-to-target retrieval tasks. Experimental results show that OpenMap outperforms state-of-the-art baselines in zero-shot settings, demonstrating the effectiveness of our method in bridging free-form language and 3D perception for embodied navigation.
Problem

Research questions and friction points this paper is trying to address.

Aligning free-form language commands with specific 3D scenes
Improving instance-level semantic consistency in visual-language mapping
Enhancing instruction interpretation for embodied navigation tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Zero-shot open-vocabulary visual-language map
Structural-Semantic Consensus constraint for aggregation
LLM-assisted Instruction-to-Instance Grounding module
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Danyang Li
Danyang Li
Shuimu Scholar, Tsinghua University
Embodied AIMobile ComputingInternet of ThingsEdge ComputingSLAM System
Z
Zenghui Yang
School of Computer Science and Engineering, Central South University
G
Guangpeng Qi
Inspur Yunzhou Industrial Internet Co., Ltd
S
Songtao Pang
Inspur Yunzhou Industrial Internet Co., Ltd
G
Guangyong Shang
Inspur Yunzhou Industrial Internet Co., Ltd
Q
Qiang Ma
Tsinghua University
Z
Zheng Yang
School of Software, Tsinghua University