Retrieve-to-Localize: Bridging Large Language Models and LiDAR Geometry for Spatial Grounding

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of precise target localization in complex spatial question answering, where existing models struggle to integrate compositional spatial reasoning with geometric grounding. We propose a novel framework that synergizes large language models (LLMs) with LiDAR geometric information. Specifically, the LLM parses spatial queries, while coordinates are derived directly from LiDAR geometry via location-aware retrieval and local point refinement mechanisms, thereby circumventing reliance on textual decoding. Furthermore, we construct the SpatialLiDAR-QA dataset to facilitate LiDAR point cloud feature alignment and conditioned proposal retrieval. Experimental results demonstrate that the proposed approach significantly outperforms mainstream LiDAR-language and multi-camera vision-language models on precise coordinate prediction tasks.
📝 Abstract
LiDAR provides precise geometric information for spatial perception tasks such as object detection in autonomous driving and outdoor robotics. However, recognizing and localizing individual objects is not sufficient to answer questions that require composing spatial relations and grounding the intended target. Motivated by recent advances in large language models (LLMs) for autonomous driving, we leverage their language priors to interpret complex spatial questions and ground the referred target in LiDAR geometry. To support this spatial grounding capability, we introduce SpatialLiDAR-QA, which combines single- and multi-step relational grounding with complementary spatial understanding tasks. We further propose SpatialLiDAR-LM, which aligns LiDAR point features with an LLM and grounds target coordinates through language-conditioned, position-aware proposal retrieval and local point refinement. This design derives target coordinates directly from local LiDAR geometry rather than through textual language decoding. Experiments demonstrate substantial improvements over representative LiDAR--language models and multi-camera VLMs on precise coordinate prediction tasks. Our dataset and model training code will be publicly released.
Problem

Research questions and friction points this paper is trying to address.

Spatial Grounding
LiDAR
Large Language Models
Autonomous Driving
Coordinate Prediction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spatial Grounding
Large Language Models
LiDAR Geometry
Retrieve-to-Localize
Point Refinement
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
B
Byounggun Park
Department of Automotive Engineering (Automotive-Computer Convergence), Hanyang University, Seoul, South Korea
G
Giyong Moon
Department of Automotive Engineering, Hanyang University, Seoul, South Korea
J
Jusung Kim
Department of Automotive Engineering (Automotive-Computer Convergence), Hanyang University, Seoul, South Korea
Soonmin Hwang
Soonmin Hwang
Department of Automotive Engineering, Hanyang University
Computer VisionAutonomous DrivingRobotics