Spatial Latent Reasoning for Embodied Reference Understanding

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of effectively organizing geometric and visual cues into intermediate supervision for visual grounding of embodied pointing gestures. To this end, we propose a spatial latent reasoning framework that introduces the first structured supervision mechanism based on ordered geometric-visual state sequences, enabling end-to-end continuous reasoning through recursive state generation. Furthermore, a parity pooling operator is incorporated to optimize feature aggregation within target regions. Combined with multimodal large model fine-tuning, the proposed method achieves up to a 21.1% improvement in mIoU on the EgoPoint-Ground dataset and surpasses existing state-of-the-art approaches by 5.2% on the YouRefIt dataset.
📝 Abstract
Pointing-gesture visual grounding requires connecting hand geometry with the visual identity and extent of a referred object. A central challenge for continuous latent reasoning is how to organize these complementary cues into useful intermediate supervision. We propose Spatial Latent Reasoning (SLR), a framework that structures this supervision around an ordered sequence of geometric and visual states. A spatial ray state is supervised by fingertip position and pointing direction, followed by four states aligned with target-region features. To construct the visual targets, we introduce parity pooling, which applies polyphase grouping to average region tokens on four interleaved spatial supports. All states are generated recurrently during training and inference; auxiliary annotations are required only during training. On EgoPoint-Ground, the framework improves mIoU over same-backbone supervised fine-tuning by 2.8, 17.5, and 21.1 percentage points on Qwen3.5-4B, Qwen2.5-VL-7B, and Qwen3-VL-8B, respectively, with improvements on both hard subsets. On YouRefIt, it achieves 77.6% precision at IoU 0.5, a numerical margin of 5.2 percentage points over the reported state of the art under differing evaluation protocols. Ablations support joint geometric and visual supervision on the standard and similar-object sets, and favor parity over three alternative pooling operators on the standard set. These results support task-structured supervision for continuous pointing grounding. We will release the code and supporting materials.
Problem

Research questions and friction points this paper is trying to address.

pointing-gesture visual grounding
embodied reference understanding
continuous latent reasoning
intermediate supervision
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spatial Latent Reasoning
Parity Pooling
Visual Grounding
Pointing Gesture
Continuous Latent Reasoning
🔎 Similar Papers
No similar papers found.