Score
Designs and implements breadth‑first search procedures that traverse a graph of semantic regions (region nodes) rather than raw low‑level states, organizing exploration frontiers and expansion order at the region level. Builds the traversal logic, loop- and revisit-avoidance mechanisms, and region-to-action mappings to produce shorter, efficient action trajectories and to prevent re-entering incorrect subregions.
This work addresses the challenge of efficiently integrating high-level natural language instructions—such as “find a cup in the kitchen”—with geometric exploration in multi-robot systems operating in unknown environments. To this end, the paper proposes the Semantic Area Graph Reasoning (SAGR) framework, which introduces, for the first time, a structured semantic area graph as an interface between large language models and multi-robot coordination. The approach constructs a semantic topological abstraction of the environment via semantic occupancy mapping, leverages a large language model for high-level semantic room assignment, and integrates deterministic frontier-based planning for local navigation. Evaluated across 100 scenes from the Habitat-Matterport3D dataset, the method achieves up to an 18.8% improvement in semantic target search efficiency while maintaining state-of-the-art exploration performance.
本文提出了一种基于语义引导的探索方法SGE,通过图像空间航点采样解决非结构化环境中的地面车辆导航问题,结合语义分割和路径优化技术。
To address two key bottlenecks in multi-robot autonomous exploration under unknown environments—insufficient spatial structural modeling of unexplored regions and high computational complexity (NP-hard) of global task allocation—this paper proposes an efficient collaborative framework based on region modeling and hierarchical planning. We introduce a novel RegionGraph weighted graph model that explicitly encodes the spatial topology and explorability of unknown regions. A hierarchical asynchronous exploration mechanism is designed to dynamically decompose global tasks into local subtasks, significantly reducing re-planning frequency. The framework integrates region segmentation, graph-based modeling, hierarchical task assignment, and frontier-point-driven exploration. Evaluated through both simulation and real-robot experiments, it achieves a 20% improvement in exploration efficiency over state-of-the-art methods, reduces mission completion time, and enhances system scalability.
This work addresses the challenges large language models face in web navigation, where sparse valid action paths and contextual noise impede accurate state perception. To overcome these limitations, the authors propose Plan-MCTS, a framework that shifts exploration into a semantic planning space, decoupling high-level strategic planning from low-level execution. The approach constructs a dense plan tree and leverages abstracted semantic history to enhance navigation efficiency. A dual-gated reward mechanism is introduced to jointly assess action executability and policy consistency, while structured refinement enables online repair of failed subplans. Evaluated on the WebArena benchmark, Plan-MCTS significantly improves both task success rates and search efficiency, achieving state-of-the-art performance.
Addressing the limitations of hand-crafted heuristics—namely, their labor-intensive design, poor scalability, and weak generalizability—in hard-exploration tasks, this paper introduces Intelligent Go-Explore (IGE). IGE is the first framework to integrate pretrained foundation models (including large language and multimodal models) into Go-Explore, leveraging their internalized “interestingness” intuition to autonomously assess state novelty and promise—eliminating reliance on human priors. IGE synergistically combines three core mechanisms: a state archive for memory retention, iterative backtracking for robust recovery, and adaptive action suggestion for directed exploration. These enable IGE to detect and exploit unforeseen yet valuable serendipitous discoveries. Empirically, IGE significantly outperforms classical reinforcement learning and graph-search baselines across diverse language- and vision-based exploration tasks. Notably, it achieves breakthrough success on tasks where state-of-the-art foundation-model agents—such as Reflexion—completely fail.
This study addresses the challenge of balancing semantic exploration with Linear Temporal Logic (LTL) task execution for robots operating in unknown environments. To this end, we propose a task-driven, non-myopic adaptive planning framework. The method leverages vision-language models to construct metric semantic scene graphs online and encodes finite LTL (LTLf) task specifications using deterministic finite automata. By incorporating real-time feedback, the framework dynamically balances exploration rewards against task progress to generate and update hybrid paths. Experimental evaluations across five task categories in photorealistic indoor environments demonstrate that, compared to baseline approaches, the proposed method achieves higher task completion rates while requiring smaller coverage areas and shorter path lengths, all while strictly satisfying task constraints.
Existing graph-based navigation methods suffer from structural error accumulation due to the absence of semantic guidance and susceptibility to erroneous predictions, compromising long-term reliability. This work proposes the Hypothesis Graph Refinement (HGR) framework, which models frontier semantic predictions as revisable hypothesis nodes and constructs a dependency-aware graph memory to support goal-directed exploration. Upon detecting conflicts between observations and predictions, HGR triggers a verification-driven cascaded error-correction mechanism that dynamically prunes incorrect subgraphs, enabling the navigation graph to contract rather than grow unidirectionally. The approach integrates vision-language models for contextual semantic prediction and introduces an exploration ranking strategy that jointly considers goal relevance, path cost, and uncertainty. Evaluated on GOAT-Bench, HGR achieves a 72.41% success rate (SPL 56.22%), substantially improving A-EQA and EM-EQA performance, while reducing redundant nodes by 20% and decreasing revisit rates in erroneous regions by 4.5×.
This work addresses the limitations of traditional autonomous drone exploration methods, which rely solely on geometric information and lack semantic awareness, resulting in inefficient exploration and insufficient environmental understanding. The study proposes a novel dynamic exploration framework that integrates semantic information into a Probabilistic Roadmap (PRM)-based approach by introducing a semantic reward function. This function guides the drone to prioritize regions containing objects and structures with high information value. An incremental update mechanism further refines frontier selection and path planning. Implemented on the ROS Noetic and Gazebo platforms, the system fuses geometric and semantic data from RGB-D sensors, incorporating real-time semantic segmentation with PRM construction. Experimental results demonstrate 90%–94% exploration coverage across diverse simulated environments, significantly outperforming conventional geometric-only methods in both exploration time and flight distance.
Existing interactive data exploration tools predominantly employ linear view sequences, limiting their ability to support branching hypothesis exploration and argument-chain construction in knowledge-intensive domains such as law. To address this, we propose SemanticTours—a nonlinear, knowledge graph–based data navigation framework designed for deep reasoning and iterative hypothesis refinement. SemanticTours introduces three core mechanisms: user-definable semantic relations, node aggregation, and semantic lensing—enabling flexible, context-aware navigation across heterogeneous legal evidence. The framework integrates knowledge graph modeling, semantic relation extraction, graph visualization, and principled interaction design, specifically optimized for complex legal case analysis. An evaluation with six domain experts demonstrates that SemanticTours’ graph-driven navigation significantly outperforms conventional linear approaches, yielding measurable improvements in analytical efficiency and expressive power of legal reasoning. These results validate both the practical utility and conceptual novelty of the proposed framework.
针对多无人机在间歇通信下探索未知环境的问题,提出了一种基于前沿连通性的新型探索图扩展策略,以提高探索效率和稳定性。