🤖 AI Summary
This study addresses the challenge of precise target localization in complex spatial question answering, where existing models struggle to integrate compositional spatial reasoning with geometric grounding. We propose a novel framework that synergizes large language models (LLMs) with LiDAR geometric information. Specifically, the LLM parses spatial queries, while coordinates are derived directly from LiDAR geometry via location-aware retrieval and local point refinement mechanisms, thereby circumventing reliance on textual decoding. Furthermore, we construct the SpatialLiDAR-QA dataset to facilitate LiDAR point cloud feature alignment and conditioned proposal retrieval. Experimental results demonstrate that the proposed approach significantly outperforms mainstream LiDAR-language and multi-camera vision-language models on precise coordinate prediction tasks.
📝 Abstract
LiDAR provides precise geometric information for spatial perception tasks such as object detection in autonomous driving and outdoor robotics. However, recognizing and localizing individual objects is not sufficient to answer questions that require composing spatial relations and grounding the intended target. Motivated by recent advances in large language models (LLMs) for autonomous driving, we leverage their language priors to interpret complex spatial questions and ground the referred target in LiDAR geometry. To support this spatial grounding capability, we introduce SpatialLiDAR-QA, which combines single- and multi-step relational grounding with complementary spatial understanding tasks. We further propose SpatialLiDAR-LM, which aligns LiDAR point features with an LLM and grounds target coordinates through language-conditioned, position-aware proposal retrieval and local point refinement. This design derives target coordinates directly from local LiDAR geometry rather than through textual language decoding. Experiments demonstrate substantial improvements over representative LiDAR--language models and multi-camera VLMs on precise coordinate prediction tasks. Our dataset and model training code will be publicly released.