Score
Design and build systems that interpret natural-language instructions and queries and align them to entities, coordinates, and regions in geometric or semantic maps. This work covers components that classify and disambiguate commands, resolve linguistic references to map elements, produce grounded prompts for vision-language models, and emit actionable targets or trajectories for downstream navigation and manipulation.
This work addresses the problem of precise grounding of natural-language queries to spatial regions in open-vocabulary vision–language maps. To this end, we propose a training-free heuristic semantic matching method. Our core innovation lies in leveraging lexical-semantic structures—specifically synonymy and antonymy relations—to construct a lightweight embedding-space guidance mechanism that jointly exploits vision–language model (VLM) features and semantic similarity metrics for efficient query-to-region mapping. The method is architecture-agnostic, supporting diverse visual encoders and text representations, thereby ensuring strong cross-domain generalization. Extensive experiments on multiple vision-map and image benchmarks demonstrate significant improvements in cross-scene query accuracy, outperforming existing baselines under zero-shot and few-shot settings. Importantly, the approach yields interpretable, low-dependency localization—requiring no large-scale supervised training—thus establishing a novel paradigm for semantic navigation and instruction-driven robotic systems.
Natural language instructions for robotic manipulation tasks often lack precise modeling of spatial relations among objects and temporal goals. Method: This paper introduces the first end-to-end framework for generating formal spatiotemporal logic (SpaTiaL) formulas directly from natural language. It (1) constructs a semantics-preserving back-translated NL2SpaTiaL dataset; (2) proposes a joint translation-verification architecture integrating a language-driven semantic validator and SMT-based satisfiability checking; and (3) synergistically combines formal modeling, deterministic back-translation, LLM fine-tuning, and logical verification. Contribution/Results: Experiments demonstrate substantial improvements in instruction grounding—enhancing interpretability, formal verifiability, and compositional generalization. Compared to conventional temporal logic (TL), spatial relation modeling coverage doubles (100% increase), and logical formula validity reaches 92.7% pass rate under SMT verification.
This work addresses the challenge of enabling non-expert users to control mobile robot navigation through natural language commands. It proposes a modular, ROS 2–based language-driven navigation framework that integrates natural language understanding, RGB-D semantic perception, and Nav2 autonomous navigation to achieve end-to-end mapping from linguistic instructions to navigational goals. The system supports context-aware instruction parsing and cross-platform deployment. Validated on both TurtleBot3 Waffle and Unitree Go2 platforms, it accurately identifies linguistically referenced targets, estimates their spatial locations, generates feasible navigation paths, and provides natural language feedback. This approach significantly enhances the intuitiveness and robustness of human–robot interaction in real-world environments.
This work investigates the capability of Large Vision-Language Models (LVLMs) to interpret pixel-level outdoor maps and generate natural-language navigation instructions. To this end, we introduce MapBench—the first outdoor navigation benchmark explicitly designed for human-readable maps—comprising 100 real-world maps and over 1,600 path-following tasks. We propose the Map Spatial Scene Graph (MSSG) as a cross-modal alignment index for fine-grained evaluation, and design a cognitively decomposed Chain-of-Thought (CoT) reasoning framework to systematically expose fundamental limitations of LVLMs in spatial reasoning and structured decision-making. Through zero-shot prompting, MSSG-guided inference, and multi-granularity evaluation, we comprehensively assess leading LVLMs, revealing an average task accuracy below 35%, confirming MapBench’s high difficulty. The benchmark dataset, evaluation code, and implementation are publicly released.
Existing robotic navigation systems struggle with unstructured, map-free environments when receiving incomplete natural-language task instructions. Method: This work proposes an online semantic planning framework leveraging large language models (LLMs), integrated into a closed-loop system that synergistically combines real-time semantic SLAM, receding-horizon planning (RHP), and online safety verification. The framework enables concurrent semantic mapping, automatic task completion, dynamic subtask re-planning, and runtime safety constraint enforcement—without requiring prior maps or manually refined commands. Contribution/Results: It establishes the first end-to-end online pipeline for semantic understanding, planning, and execution. Evaluated in cluttered outdoor environments exceeding 20,000 m², the approach reduces task completion time and path length by over 50% compared to baselines, while significantly decreasing user interaction frequency.
This work proposes a semantic navigation approach for autonomous robots operating in unknown environments, addressing the common oversight of semantic cues such as signs and room numbers that can significantly improve navigation efficiency. The method integrates local perception, frontier-based exploration, and a large language model (LLM) to dynamically interpret environmental text, infer symbolic patterns (e.g., room numbering conventions), and construct a confidence grid that guides exploration in real time. Notably, this is the first framework to employ an LLM for on-the-fly parsing of environmental text to drive forward-looking goal inference. Evaluated on realistic floorplans, the approach achieves a weighted path success rate over 25% higher than baseline methods, approaching the performance of optimal paths.
Large language models lack native support for continuous spatial representations, hindering their capacity for genuine geometric reasoning. This work proposes the Spatial Language Model (SLM), which for the first time integrates learnable spatial representations as a first-class modality directly into the model’s reasoning process. By leveraging a multimodal architecture, atomic geometric operations, and training with spatial instruction alignment, SLM achieves a paradigm shift from symbolic matching to true geometric reasoning. Accompanied by a newly curated spatial instruction dataset and the SpatialEval benchmark, empirical evaluations demonstrate that SLM substantially outperforms existing approaches—whether based on prompt engineering or textual abstraction—across tasks involving spatial attributes, distance, topology, and relative positioning.
This study addresses the challenges of high noise levels and sparse semantics in raw vessel AIS trajectory data by proposing a context-aware trajectory abstraction framework. The method segments trajectories into voyages and movement segments, integrates multi-source contextual information—including geographic entities, navigational characteristics, and weather conditions—and leverages semantic annotation combined with a large language model (LLM)-guided generation mechanism. This approach enables, for the first time, the production of natural language descriptions with high semantic density driven by heterogeneous contextual cues. The resulting structured and interpretable trajectory representations substantially reduce spatiotemporal complexity, thereby enhancing the usability of maritime data and supporting more effective high-level reasoning.
To address the challenge of precisely grounding open-vocabulary, free-form natural language instructions to 3D scene instances in open-world embodied agents, this paper proposes a zero-shot open-vocabulary vision-language grounding method. Methodologically, it integrates (1) a structure-semantic consistency constraint to ensure geometric and semantic alignment of multi-view instance representations; (2) an LLM-assisted instruction-to-instance alignment module for fine-grained semantic understanding; and (3) a unified framework combining vision-language models, 3D instance aggregation, explicit geometric modeling, and LLM-driven spatial contextual reasoning. Experiments on ScanNet200 and Matterport3D demonstrate substantial improvements over state-of-the-art methods in semantic mapping and instruction-based target retrieval. Notably, our approach achieves, for the first time, open-vocabulary, instance-level instruction grounding without requiring category priors or labeled training data.