Soft Spatial Reasoning

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the premature discretization and error propagation issues caused by hard thinking in the spatial reasoning of large vision-language models (LVLMs) by proposing a soft spatial reasoning framework. This approach constructs continuous soft states through mixed word embeddings, enabling multiple candidate paths to collaboratively influence subsequent reasoning. Its core innovation, the AdaptSoft controller, leverages hidden states and predictive uncertainty to adaptively modulate the degree of softening at each step. Combined with post-training techniques and a gradient alignment objective, the framework can be optimized without intermediate supervision signals. Experimental results demonstrate that it significantly outperforms hard thinking, fixed-softening baselines, and mainstream LVLMs across multiple spatial reasoning benchmarks.
📝 Abstract
Large Vision-Language Models (LVLMs) commonly perform spatial reasoning through chain-of-thought (CoT), encoding intermediate reasoning as autoregressive sequences of discrete language tokens. Such hard thinking requires committing to a single token at each step, even when the correct spatial interpretation remains uncertain. This early commitment constitutes premature discretization: an incorrect token selection can propagate errors through subsequent reasoning. We propose Soft Spatial Reasoning, a post-training framework that introduces soft thinking for spatial tasks in LVLMs. At each intermediate reasoning step, the LVLM forms a continuous soft state by mixing token embeddings rather than selecting a single token, allowing multiple candidate continuations to influence the next step. The appropriate degree of softness, however, can vary across reasoning steps: retaining multiple candidates may preserve a useful spatial interpretation, but if those candidates imply conflicting spatial relations, mixing them may interfere with subsequent reasoning. At the core of Soft Spatial Reasoning is AdaptSoft, a controller that uses the current hidden state and predictive uncertainty to adapt the degree of softness at each reasoning step. To train AdaptSoft, we introduce a gradient-alignment learning objective that provides a step-specific learning signal for softness control without intermediate reasoning supervision. Across diverse spatial benchmarks, Soft Spatial Reasoning outperforms hard and fixed-soft CoT baselines using the same backbone, as well as a range of existing LVLMs. The source code is available at https://github.com/rafiibnsultan/Soft_Spatial_Reasoning
Problem

Research questions and friction points this paper is trying to address.

Spatial Reasoning
Large Vision-Language Models
Premature Discretization
Chain-of-Thought
Error Propagation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Soft Spatial Reasoning
Premature Discretization
AdaptSoft
Gradient-Alignment Learning
Vision-Language Models
🔎 Similar Papers
No similar papers found.