SpatiaLoc: Leveraging Multi-Level Spatial Enhanced Descriptors for Cross-Modal Localization

πŸ“… 2026-01-07
πŸ›οΈ arXiv.org
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the problem of text-to-point-cloud cross-modal localization, aiming to enable precise 2D robot localization within an environment using natural language descriptions. The proposed approach employs a coarse-to-fine strategy: in the coarse localization stage, a BΓ©zier-enhanced object spatial encoder and a frequency-domain-aware encoder are introduced to jointly model instance-level and global-level spatial relationships; in the fine localization stage, an uncertainty-aware Gaussian regression mechanism is incorporated to refine positional accuracy. This framework is the first to integrate multi-granularity spatial relationship modeling into cross-modal localization and achieves state-of-the-art performance on the KITTI360Pose dataset, demonstrating significantly improved accuracy and robustness over existing methods.

Technology Category

Intelligent Robots: Localization, Mapping, and NavigationNatural Language Processing: Language Grounding & Multi-modal NLPComputer Vision: Multi-modal Vision

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSystems and Infrastructure for Web, Mobile and WoT: Location- and context-aware Web and WoT applications and services
πŸ“ Abstract
Cross-modal localization using text and point clouds enables robots to localize themselves via natural language descriptions, with applications in autonomous navigation and interaction between humans and robots. In this task, objects often recur across text and point clouds, making spatial relationships the most discriminative cues for localization. Given this characteristic, we present SpatiaLoc, a framework utilizing a coarse-to-fine strategy that emphasizes spatial relationships at both the instance and global levels. In the coarse stage, we introduce a Bezier Enhanced Object Spatial Encoder (BEOSE) that models spatial relationships at the instance level using quadratic Bezier curves. Additionally, a Frequency Aware Encoder (FAE) generates spatial representations in the frequency domain at the global level. In the fine stage, an Uncertainty Aware Gaussian Fine Localizer (UGFL) regresses 2D positions by modeling predictions as Gaussian distributions with a loss function aware of uncertainty. Extensive experiments on KITTI360Pose demonstrate that SpatiaLoc significantly outperforms existing state-of-the-art (SOTA) methods.
Problem

Research questions and friction points this paper is trying to address.

cross-modal localization
spatial relationships
text-to-point cloud
robot navigation
natural language grounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

cross-modal localization
spatial relationship modeling
Bezier curve encoding
frequency domain representation
uncertainty-aware localization
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
T
Tianyi Shang
Fuzhou University
P
Pengjie Xu
Qingdao University
Z
Zhaojun Deng
Tongji University
Zhenyu Li
Zhenyu Li
Qilu University of Technology (Shandong Academy of Sciences)
Intelligent PerceptionComputer VisionPlace RecognitionEdge Computing
Z
Zhicong Chen
Fuzhou University
Lijun Wu
Lijun Wu
Shanghai AI Laboratory
MLLLMAI4Science