🤖 AI Summary
This study addresses the instability of existing 3D large models in complex scene reasoning, which stems from spatial relationship loss during passive compression. To overcome this limitation, this work proposes an active scene state framework that shifts the paradigm from passive feature compression to proactive scene organization. Methodologically, it decouples entity, frame, relational, and global representations through superpoint-level evidence aggregation and multi-granularity modeling. Furthermore, a spatial anchoring mechanism is introduced to generate role-aware scaffolds, enabling collaborative reasoning with large language models. The proposed approach yields substantial performance improvements across 3D visual grounding, question answering, and dense captioning tasks, effectively resolving challenges associated with spatial ambiguity and relation-dense scenarios.
📝 Abstract
Recent 3D large multimodal models (3D-LMMs) rely on a visual bottleneck to compress complex 3D scene evidence into a limited number of visual tokens compatible with large language models (LLMs). Current visual bottlenecks, however, often passively compress heterogeneous 3D evidence into a homogeneous object-centric token sequence, leaving the spatial organization of the scene under-represented. This under-representation forces the LLM to recover spatial relations from a flattened token sequence, leading to unstable reasoning in relation-intensive and spatially ambiguous scenes. To address this issue, we propose SceneScaffold, an active scene-state construction framework for unified 3D scene understanding. SceneScaffold reformulates the visual bottleneck from a passive feature compressor into an active scene organizer, constructing a role-aware spatial scaffold before language reasoning. Specifically, SceneScaffold organizes superpoint-level visual evidence into scene-state components with distinct structural roles: entity states preserve core object semantics, scene-frame states maintain spatial references via boundary and region anchors, relation states encode object-environment interaction cues, and a global summary provides compact context. Through this role-aware construction, SceneScaffold provides the LLM with a spatially organized scene representation before language reasoning. Experiments on unified 3D scene understanding tasks, including 3D visual grounding, question answering, and dense captioning, demonstrate the effectiveness of SceneScaffold, while diagnostic results further show its applicability to relation-intensive and spatially ambiguous cases. Code is available at https://github.com/lixiangqi707/SceneScaffold.