SpatialCORE: Confidence-Aware Grounded Spatial Reasoning in Large Vision--Language Models

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of large vision-language models in spatial reasoning, where the absence of localization confidence constraints often leads to correct answers accompanied by flawed reasoning. To mitigate this issue, this work proposes SpatialCORE, a novel framework that pioneers the integration of coordinate token confidence into spatial reward mechanisms by leveraging the model's intrinsic grounding confidence as a learning signal. Through post-training with self-regulating rewards and an answer-gating mechanism, the framework achieves joint optimization of grounding quality and reasoning correctness. Extensive experiments demonstrate that SpatialCORE attains state-of-the-art performance across multiple benchmarks among both open-source and specialized spatial reasoning models. Furthermore, it exhibits robust generalization to unseen data distributions under zero-shot settings.
📝 Abstract
Large Vision-Language Models (LVLMs) have made remarkable progress across visual perception tasks, yet spatial reasoning remains a persistent weakness, especially for questions that require reasoning over visual space. Recent spatial-reasoning methods incorporate generated grounding, where models predict bounding boxes, masks, or other localization outputs for task-relevant objects as part of their reasoning trace. However, these approaches typically optimize final-answer correctness alone, allowing correct answers to be rewarded even when the model does not reason from confidently localized task-relevant objects. We introduce SpatialCORE (Spatially COnfident REasoning), a post-training framework that turns the model's own confidence in generated grounding into a learning signal for spatial reasoning. Its central idea is to reinforce grounding that is both accurate and confident, encouraging the model to reason from confidently localized task-relevant objects. SpatialCORE realizes this through a self-regulating spatial reward that weights each predicted bounding box's matching quality by its coordinate-token confidence. An answer gate further ties grounding optimization to final-answer correctness. SpatialCORE achieves state-of-the-art results among open-source and specialized spatial reasoning models across diverse benchmarks, and transfers effectively in zero-shot settings to unseen data distributions. The source code is available at https://github.com/rafiibnsultan/SpatialCORE.
Problem

Research questions and friction points this paper is trying to address.

Spatial Reasoning
Large Vision-Language Models
Grounding
Confidence-Aware
Visual Perception
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spatial Reasoning
Confidence-Aware Grounding
Self-Regulating Spatial Reward
Large Vision-Language Models
Answer Gate
🔎 Similar Papers
No similar papers found.