LOCUS: Landmark-Oriented Container Discrimination Using Spatial Graphs

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of localizing hidden objects in unstructured environments, where robots struggle to distinguish between spatially and semantically similar items. To overcome this, we propose a joint spatial-semantic reasoning framework based on scene graphs. Specifically, a Graph Neural Network (GNN) updates CLIP embeddings according to object proximity to enhance container identification. Furthermore, this work pioneers the integration of household ontological knowledge into CLIP representations, achieving spatial awareness enhancement through message passing among graph nodes. The proposed system synergistically combines Detic detection, CLIP multimodal modeling, and GNN techniques. Experimental results demonstrate that our approach surpasses existing baselines in simulation and exhibits robustness against label noise during real-world exploration tasks conducted with a mobile manipulator.
📝 Abstract
As robots are increasingly deployed in unstructured, real-world environments, the ability to reason about complex spatial and semantic relationships among objects remains a fundamental challenge in enabling robust and generalizable manipulation and navigation. For example, deciding where to search for an object that is not in plain sight depends on where it is physically plausible as well as semantically likely. Popular methods such as cosine similarity with CLIP struggle to disambiguate between spatially and semantically similar objects that could contain a target object. We propose Landmark-Oriented Container Discrimination Using Spatial Graphs (LOCUS). To jointly reason about spatial information and semantics, we train a GNN to update CLIP embeddings on a full environment scene graph by passing embedding information between nodes based on proximity. CLIP embeddings, augmented by semantic knowledge from a fusion of household ontologies provide a robust semantic signal. We evaluate in simulation on a noiseless scene graph from simulator metadata, making the approach detection-agnostic. Our approach outperforms random, CLIP, Tidybot, and an LLM planner in the majority of room classes and configurations in simulation. To show our approach is robust to scene graph node placement and label noise, we demonstrate a physical mobile manipulator running in an exploration pipeline start-to-finish including scene graph labels from Detic.
Problem

Research questions and friction points this paper is trying to address.

spatial reasoning
semantic reasoning
object search
container discrimination
robot navigation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spatial Graphs
Graph Neural Network
CLIP Embeddings
Scene Graph
Container Discrimination
T
Taylor Bergeron
Department of Robotics Engineering, Worcester Polytechnic Institute, Worcester, MA, USA
S
Shibani Senthilbabu
Department of Robotics Engineering, Worcester Polytechnic Institute, Worcester, MA, USA
Kevin Leahy
Kevin Leahy
Worcester Polytechnic Institute
RoboticsFormal MethodsMulti-Agent Systems