🤖 AI Summary
This study addresses a critical limitation in multimodal large language models: while they can accurately report spatial facts, they struggle to effectively leverage this information during subsequent reasoning. We identify this "availability-utilization" gap and construct the SpaceConflict benchmark along with a unified judgment interface to quantify it. To bridge this divide, we propose Operational State Supervision (OSS), which explicitly supervises spatial states and their transformation trajectories to align context-shared facts. Experimental results demonstrate that OSS significantly improves L3/L4-level pairwise accuracy, confirming that scaling alone cannot close the state utilization gap. This work establishes a new paradigm for enhancing the spatial reasoning capabilities of multimodal models.
📝 Abstract
A multimodal large language model that correctly reports a spatial fact does not necessarily use that fact in subsequent reasoning. To study this distinction, we introduce \textsc{SpaceConflict}, a benchmark of 23{,}196 inputs for the construction and use of spatial state. Under a unified Supported/Contradictory/Unknown judgment interface, it covers local fact binding (L1), relational composition (L2), cross-observation consistency (L3), and state judgment under transformation (L4). Posing a direct-state query, a full-transformation query, and an explicit-initial-state query on the same world reveals an availability--utilization gap: models recover the initial state from visual evidence yet fail when that state must drive a transformation. For Qwen3.5-9B, 50 of 100 sequences with a correctly recovered initial state fail the full transformation, and supplying the state explicitly repairs all 50; the gap narrows with scale but does not close. We therefore propose Operational State Supervision (OSS), which supervises task-relevant spatial states and their transformation trajectories and aligns shared facts across contexts. OSS improves paired accuracy on matched judgments most on L3 and L4, the levels that depend on organizing and using state. Evaluating multimodal spatial reasoning thus requires asking not only whether a model can see and state a spatial fact, but whether that fact becomes a usable state in subsequent computation.