language-to-map grounding

Design and build systems that interpret natural-language instructions and queries and align them to entities, coordinates, and regions in geometric or semantic maps. This work covers components that classify and disambiguate commands, resolve linguistic references to map elements, produce grounded prompts for vision-language models, and emit actionable targets or trajectories for downstream navigation and manipulation.

language-to-mapgrounding

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.45
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

QuASH: Using Natural-Language Heuristics to Query Visual-Language Robotic Maps

Oct 16, 2025
MP
Matti Pekkanen
🏛️ Aalto University

This work addresses the problem of precise grounding of natural-language queries to spatial regions in open-vocabulary vision–language maps. To this end, we propose a training-free heuristic semantic matching method. Our core innovation lies in leveraging lexical-semantic structures—specifically synonymy and antonymy relations—to construct a lightweight embedding-space guidance mechanism that jointly exploits vision–language model (VLM) features and semantic similarity metrics for efficient query-to-region mapping. The method is architecture-agnostic, supporting diverse visual encoders and text representations, thereby ensuring strong cross-domain generalization. Extensive experiments on multiple vision-map and image benchmarks demonstrate significant improvements in cross-scene query accuracy, outperforming existing baselines under zero-shot and few-shot settings. Importantly, the approach yields interpretable, low-dependency localization—requiring no large-scale supervised training—thus establishing a novel paradigm for semantic navigation and instruction-driven robotic systems.

Identifying environment parts relevant to natural language queriesLeveraging synonyms and antonyms in embedding spaceTraining classifiers to partition environments for query matching

NL2SpaTiaL: Generating Geometric Spatio-Temporal Logic Specifications from Natural Language for Manipulation Tasks

Dec 15, 2025
LL
Licheng Luo
🏛️ University of California, Riverside | Lehigh University

Natural language instructions for robotic manipulation tasks often lack precise modeling of spatial relations among objects and temporal goals. Method: This paper introduces the first end-to-end framework for generating formal spatiotemporal logic (SpaTiaL) formulas directly from natural language. It (1) constructs a semantics-preserving back-translated NL2SpaTiaL dataset; (2) proposes a joint translation-verification architecture integrating a language-driven semantic validator and SMT-based satisfiability checking; and (3) synergistically combines formal modeling, deterministic back-translation, LLM fine-tuning, and logical verification. Contribution/Results: Experiments demonstrate substantial improvements in instruction grounding—enhancing interpretability, formal verifiability, and compositional generalization. Compared to conventional temporal logic (TL), spatial relation modeling coverage doubles (100% increase), and logical formula validity reaches 92.7% pass rate under SMT verification.

Addresses lack of datasets capturing spatial relations in manipulation instructionsEnsures semantic fidelity between language and logic via verification frameworkGenerates SpaTiaL logic from natural language for robotic manipulation tasks

This work addresses the challenge of enabling non-expert users to control mobile robot navigation through natural language commands. It proposes a modular, ROS 2–based language-driven navigation framework that integrates natural language understanding, RGB-D semantic perception, and Nav2 autonomous navigation to achieve end-to-end mapping from linguistic instructions to navigational goals. The system supports context-aware instruction parsing and cross-platform deployment. Validated on both TurtleBot3 Waffle and Unitree Go2 platforms, it accurately identifies linguistically referenced targets, estimates their spatial locations, generates feasible navigation paths, and provides natural language feedback. This approach significantly enhances the intuitiveness and robustness of human–robot interaction in real-world environments.

mobile robotsnatural language interactionnavigation goals

Can Large Vision Language Models Read Maps Like a Human?

Mar 18, 2025
SX
Shuo Xing
🏛️ Texas A&M University | MBZUAI | UC Berkeley | University of Michigan | UC Riverside

This work investigates the capability of Large Vision-Language Models (LVLMs) to interpret pixel-level outdoor maps and generate natural-language navigation instructions. To this end, we introduce MapBench—the first outdoor navigation benchmark explicitly designed for human-readable maps—comprising 100 real-world maps and over 1,600 path-following tasks. We propose the Map Spatial Scene Graph (MSSG) as a cross-modal alignment index for fine-grained evaluation, and design a cognitively decomposed Chain-of-Thought (CoT) reasoning framework to systematically expose fundamental limitations of LVLMs in spatial reasoning and structured decision-making. Through zero-shot prompting, MSSG-guided inference, and multi-granularity evaluation, we comprehensively assess leading LVLMs, revealing an average task accuracy below 35%, confirming MapBench’s high difficulty. The benchmark dataset, evaluation code, and implementation are publicly released.

Assesses spatial reasoning and decision-making in LVLMsEvaluates LVLMs on human-readable map navigation tasksIntroduces MapBench dataset for complex outdoor navigation scenarios

Existing robotic navigation systems struggle with unstructured, map-free environments when receiving incomplete natural-language task instructions. Method: This work proposes an online semantic planning framework leveraging large language models (LLMs), integrated into a closed-loop system that synergistically combines real-time semantic SLAM, receding-horizon planning (RHP), and online safety verification. The framework enables concurrent semantic mapping, automatic task completion, dynamic subtask re-planning, and runtime safety constraint enforcement—without requiring prior maps or manually refined commands. Contribution/Results: It establishes the first end-to-end online pipeline for semantic understanding, planning, and execution. Evaluated in cluttered outdoor environments exceeding 20,000 m², the approach reduces task completion time and path length by over 50% compared to baselines, while significantly decreasing user interaction frequency.

Enhancing efficiency and reducing user interaction in robotic missionsMapping and planning in unstructured environments without pre-built mapsOnline semantic planning for incomplete natural language missions

Latest Papers

What's happening recently
View more

This work proposes a semantic navigation approach for autonomous robots operating in unknown environments, addressing the common oversight of semantic cues such as signs and room numbers that can significantly improve navigation efficiency. The method integrates local perception, frontier-based exploration, and a large language model (LLM) to dynamically interpret environmental text, infer symbolic patterns (e.g., room numbering conventions), and construct a confidence grid that guides exploration in real time. Notably, this is the first framework to employ an LLM for on-the-fly parsing of environmental text to drive forward-looking goal inference. Evaluated on realistic floorplans, the approach achieves a weighted path success rate over 25% higher than baseline methods, approaching the performance of optimal paths.

autonomous navigationgoal-directed navigationsemantic navigation

Large language models lack native support for continuous spatial representations, hindering their capacity for genuine geometric reasoning. This work proposes the Spatial Language Model (SLM), which for the first time integrates learnable spatial representations as a first-class modality directly into the model’s reasoning process. By leveraging a multimodal architecture, atomic geometric operations, and training with spatial instruction alignment, SLM achieves a paradigm shift from symbolic matching to true geometric reasoning. Accompanied by a newly curated spatial instruction dataset and the SpatialEval benchmark, empirical evaluations demonstrate that SLM substantially outperforms existing approaches—whether based on prompt engineering or textual abstraction—across tasks involving spatial attributes, distance, topology, and relative positioning.

continuous spatial representationsgeometric reasoninglarge language models

This study addresses the challenges of high noise levels and sparse semantics in raw vessel AIS trajectory data by proposing a context-aware trajectory abstraction framework. The method segments trajectories into voyages and movement segments, integrates multi-source contextual information—including geographic entities, navigational characteristics, and weather conditions—and leverages semantic annotation combined with a large language model (LLM)-guided generation mechanism. This approach enables, for the first time, the production of natural language descriptions with high semantic density driven by heterogeneous contextual cues. The resulting structured and interpretable trajectory representations substantially reduce spatiotemporal complexity, thereby enhancing the usability of maritime data and supporting more effective high-level reasoning.

AIS datacontext enrichmentnatural language description

OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping

Aug 03, 2025
DL
Danyang Li
🏛️ Tsinghua University | Central South University | Inspur Yunzhou Industrial Internet Co., Ltd

To address the challenge of precisely grounding open-vocabulary, free-form natural language instructions to 3D scene instances in open-world embodied agents, this paper proposes a zero-shot open-vocabulary vision-language grounding method. Methodologically, it integrates (1) a structure-semantic consistency constraint to ensure geometric and semantic alignment of multi-view instance representations; (2) an LLM-assisted instruction-to-instance alignment module for fine-grained semantic understanding; and (3) a unified framework combining vision-language models, 3D instance aggregation, explicit geometric modeling, and LLM-driven spatial contextual reasoning. Experiments on ScanNet200 and Matterport3D demonstrate substantial improvements over state-of-the-art methods in semantic mapping and instruction-based target retrieval. Notably, our approach achieves, for the first time, open-vocabulary, instance-level instruction grounding without requiring category priors or labeled training data.

Aligning free-form language commands with specific 3D scenesEnhancing instruction interpretation for embodied navigation tasksImproving instance-level semantic consistency in visual-language mapping

Hot Scholars

WZ

Wei Zhan

Co-Director of Berkeley DeepDrive, UC Berkeley; Chief Scientist of Applied Intuition
AI for autonomous systems
XH

Xiao He

AI2Robotics
computer vision
HX

Hongyu Xu

Research Scientist, Meta Reality Labs
Spatial PerceptionGenAIMultimodalRoomPlan
MD

Mingyu Ding

Assistant Professor, UNC Chapel Hill
RoboticsEmbodied AIComputer Vision
JG

Jingyu Gong

Shanghai Jiao Tong University
3D Computer Vision