GeoLatent: Geometry-Guided Latent Structuring with Routed Optimization for 3D Reasoning

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of 3D spatial reasoning from 2D images and the geometric collapse of continuous latent variables by proposing the GeoLatent framework. Specifically, it introduces a Common-Residual Geometric Alignment (CR-GEO) mechanism that disentangles shared and residual geometric features to effectively prevent geometric directional collapse. Furthermore, a routing optimization strategy is designed to enforce visual answer generation through latent variable mediation, thereby enhancing the effective rank of representations. By integrating decomposed spatial latents with multi-task joint training, the framework achieves structured geometric representations. Experimental results demonstrate that the model increases the geometric effective rank to 3.87 and attains accuracies of 73.0% and 72.1% on SPAR-Bench and SPBench, respectively, significantly outperforming existing methods.
📝 Abstract
Despite progress in vision-language models, 3D spatial reasoning from 2D images remains challenging. Text-based methods describe intermediate geometry with discrete tokens, limiting fidelity for continuous spatial relations. Continuous latents offer richer representations, but a single latent type does not explicitly separate the cues needed across spatial tasks. Decomposed spatial latents address this by representing position, direction, and global geometry separately under geometric supervision. Yet the geometry representation can still collapse toward one dominant direction, and unrestricted attention can leave the latents underused during answer learning. We introduce GeoLatent, combining Common--Residual Geometry Alignment (CR-GEO) with routed optimization to structure the geometry states while promoting latent-mediated answer learning. CR-GEO separates shared from residual teacher geometry; routed optimization jointly trains geometry and language, temporarily directs visual answer learning through the latents, and restores full attention with geometry supervision. In controlled comparisons, CR-GEO raises geometry effective rank from 1.00 to 3.87, while blocking latent readout at the bottleneck lowers direction accuracy from 89.1% to 25.8% on 128 fixed questions. After recovery, the differentiated geometry representation and latent-mediated visual route remain available alongside direct image access. GeoLatent achieves 73.0% on SPAR-Bench and 72.1% on SPBench, outperforming previously reported methods on both.
Problem

Research questions and friction points this paper is trying to address.

3D spatial reasoning
vision-language models
latent representation
geometry collapse
attention optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

3D spatial reasoning
Geometry-guided latent structuring
Common-Residual Geometry Alignment (CR-GEO)
Routed optimization
Vision-language models
🔎 Similar Papers