Toward Comprehensive 3D Grounding: Orientation Grounding through Vision-Language Models

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing 3D vision-language models (VLMs), which support only object localization and lack explicit understanding of orientation and symmetry. To bridge this gap, we introduce a novel task termed "orientation grounding," construct the ReferOri dataset, and propose the OG-VLM framework. Our method incorporates structured outputs, symmetry tokens, and geometry-aware auxiliary losses to enable joint prediction of 6D orientations and axial symmetries through multi-view reconstruction and consistency verification. Extensive experiments demonstrate that OG-VLM significantly outperforms existing baselines and object-level foundation models on both single- and multi-view benchmarks. These findings validate the learnability of orientation grounding and extend the geometric perception capabilities of 3D VLMs beyond conventional spatial localization.
📝 Abstract
Grounding is a core capability of spatial vision-language models, yet most existing work focuses only on where a referred object is. Many 3D tasks also require knowing how it is oriented. Although existing 3D VLMs may predict oriented boxes, box pose does not explicitly capture object-centric orientation or symmetry-induced ambiguities. We introduce orientation grounding, a referring grounding task that predicts an object's 6D orientation and axial symmetry from a language or box query in single-view or multi-view scenes. To support this task, we construct ReferOri, with 331K multi-view and 387K single-view orientation-grounding queries obtained through scalable reconstruction, consistency checking, and human verification. We further present OG-VLM, which adapts a 3D VLM with structured box/orientation outputs, sign and symmetry tokens, and geometry-aware auxiliary losses. Across single-view and multi-view benchmarks, OG-VLM substantially outperforms orientation-aware VLM baselines and surpasses object-level orientation foundation models on scene-level referring benchmarks, showing that explicit orientation grounding is a distinct and learnable capability beyond localization. Downstream results validate its benefit for orientation-related spatial reasoning.
Problem

Research questions and friction points this paper is trying to address.

3D grounding
orientation grounding
vision-language models
6D orientation
axial symmetry
Innovation

Methods, ideas, or system contributions that make the work stand out.

Orientation Grounding
6D Orientation
Vision-Language Models
ReferOri
OG-VLM
🔎 Similar Papers
No similar papers found.