Score
Designs, builds, or evaluates algorithms and models that infer the composition, geometry, semantics, and spatial relationships of a scene from visual or sensor inputs; produces structured scene representations such as semantic and instance labels, depth or geometry maps, object poses/bounding volumes, and scene graphs for use by downstream perception, reasoning, or interaction components.
Industrial CAD models often lack semantic, spatial, and functional information, limiting their utility in robotic simulation and high-level scene understanding. This work proposes an offline method that leverages large vision-language models (LVLMs)—introduced for the first time into CAD environments—to automatically generate structured 3D scene graphs that explicitly model manipulable objects and their functional relationships. By effectively integrating semantic parsing with functional reasoning, the approach achieves high-precision semantic annotation and relationship recognition on industrial structures such as piping systems. Both qualitative and quantitative evaluations demonstrate its effectiveness. The associated code and dataset have been made publicly available.
This study systematically evaluates the representational capabilities of vision-language models (VLMs) and video generation models (VGMs) on spatial intelligence tasks. Using a frozen-feature probing approach, the analysis compares their performance across three dimensions: semantic labeling, instance grouping, and 3D geometric prediction. The work reveals, for the first time, a complementary relationship between VLMs and VGMs in spatial understanding: VLMs excel at semantic and instance-level recognition, whereas VGMs demonstrate superior modeling of geometric structure and camera motion dynamics. Notably, a simple fusion of their representations yields substantial gains in overall performance, simultaneously enhancing both semantic accuracy and geometric fidelity.
This study investigates the capabilities and limitations of multimodal generative models (e.g., DALL-E 3, GPT-4V) in compositional scene understanding—specifically, scenes involving more than five objects and multiple spatial or semantic relations—and quantifies their performance gap relative to human cognition. To this end, we introduce the first standardized benchmark for compositional visual reasoning, comprising a structured test set generated via controllable synthetic prompting, a unified cross-model evaluation protocol, and a human–model comparative experimental paradigm. Results show that while current models significantly outperform prior generations on simple compositional tasks, their accuracy drops sharply on complex scenes, averaging 42.6% lower than human performance—revealing a fundamental bottleneck in structured visual reasoning. Our core contribution is the first comprehensive evaluation framework for compositional scene understanding with an explicit human baseline, which clearly exposes the generational gap between state-of-the-art models and humans in symbolic spatial relation modeling.
This work addresses the challenge of simultaneously achieving structural precision, semantic interpretability, and identity controllability in existing 3D/4D scene representations. We propose “Scene Language”—a unified 3D/4D scene representation framework that integrates executable program structures, natural-language semantic tokens, and visual identity embeddings. To our knowledge, this is the first method enabling zero-shot cross-modal reasoning: without fine-tuning, it directly synthesizes structured scene programs from pretrained language models and vision encoders, while explicitly modeling hierarchical relationships to support fine-grained editing. The representation is renderer-agnostic, interfacing seamlessly with traditional, neural, and hybrid renderers to produce high-fidelity images. Experiments demonstrate significant improvements over baselines—including scene graphs—on complex scene generation tasks, achieving breakthroughs in fidelity, controllability, and editability.
This work investigates the capacity of large language models (LLMs) to perform spatial semantic understanding and cross-modal reasoning solely from symbolic graphical programs—such as curve parameters, stroke sequences, and local curvature—without visual encoders. To this end, we introduce the first benchmark for symbolic graphical program-based visual understanding, comprising three tasks: program generation, semantic question answering, and zero-visual-input cross-modal reasoning. We propose Symbolic Instruction Tuning (SIT), a novel fine-tuning paradigm that explicitly enhances LLMs’ spatial reasoning capabilities using synthetic symbolic graphical instruction data. Experimental results demonstrate that strong reasoning-oriented LLMs achieve superior performance; SIT substantially improves accuracy on symbolic graphical understanding tasks and—unexpectedly—generalizes to multiple general-purpose reasoning benchmarks (e.g., GSM8K, MMLU), indicating that symbolic spatial representations can strengthen foundational reasoning abilities.
This work addresses the insufficient integration of learning-based methods and geometric constraints in camera pose and scene structure estimation by proposing a modular framework. The approach first employs a learning model (VGGT) to generate initial hypotheses for depth and relative pose, which are subsequently refined and validated using classical geometric algorithms such as point-to-plane RGB-D ICP. Crucially, the framework explicitly distinguishes the roles of learning as a “proposer” and geometry as a “referee,” emphasizing that the geometric module serves not merely as post-processing but as an essential mechanism for verifying and integrating learned outputs. Experiments on the TUM RGB-D dataset demonstrate that, in moderately challenging rigid scenes, the system significantly outperforms both purely learning-based and purely geometric baselines when the learned depth aligns geometrically with the camera intrinsics and undergoes optimization by the geometric backend.
This work addresses the limitations of existing 3D scene graph methods, which are constrained by predefined relationship categories and struggle to capture open-ended semantics and causal connections. To overcome this, the authors propose a novel framework that integrates vision-language models (VLMs) with large language models (LLMs) to construct a hierarchical forest of 3D semantic scene graphs. The VLM extracts instance-level nodes and geometry-aware relationships, while the LLM performs high-level reasoning to generate abstract concepts and open-vocabulary semantic associations. This approach transcends the confines of closed relationship sets, substantially enhancing the semantic depth and expressiveness of scene representations. Experiments on uHumans2 and ScanNet demonstrate improved accuracy in relationship generation, and real-world deployment on a Spot robot successfully enables open-vocabulary object retrieval in physical environments.
Existing approaches to 3D scene understanding are largely confined to semantic information, often neglecting the modeling of physical properties and articulated structures, and exhibit limited generalization. This work proposes the first unified framework that integrates symbolic reasoning with structured 3D geometry to reconstruct object-centric 3D representations from RGB-D observations. The method associates object instances across viewpoints, decomposes them into functional parts, and jointly infers material properties and articulation parameters, thereby constructing a scene graph that is both semantically meaningful and physically consistent. Evaluated on both synthetic and real-world datasets, the approach achieves state-of-the-art performance in semantic segmentation, multi-object centroid estimation, and joint articulation prediction. Furthermore, it demonstrates successful application in constraint-aware 3D affordance prediction and real-to-sim transfer tasks.
This work addresses the challenge that multimodal large language models often struggle to accurately localize targets in visually dense tasks due to their neglect of structured relationships among objects. To overcome this limitation, the authors propose the “Scene Graph Reasoning” (SaGe) paradigm, which explicitly integrates scene graphs into multimodal reasoning for the first time. An automated data engine constructs hierarchical scene graphs with relational edges from image–text corpora and samples 120,000 structured reasoning trajectories. The model is then trained via a two-stage graph alignment strategy: first, supervised fine-tuning internalizes structured reasoning capabilities, followed by reinforcement fine-tuning that introduces a node-agent reward mechanism to optimize graph exploration efficiency. The approach achieves significant performance gains across eight mainstream benchmarks, particularly excelling in fine-grained visual perception and reasoning tasks.
This work addresses the limited spatial understanding and layout consistency of current large language models and vision-language models in fine-grained visual editing. To overcome this, the authors propose a structured reasoning framework that explicitly models scene graph relationships to enable controllable and interpretable spatial layout editing guided by natural language instructions. The approach integrates scene graph representations, structured relational reasoning, and language guidance within a contrastive training paradigm, surpassing the limitations of conventional end-to-end methods and chain-of-thought supervised fine-tuning or GRPO strategies. Evaluated on a newly introduced benchmark for text-guided layout editing, the method achieves a 15% improvement in average IoU, reduces center distance error by 25%, and outperforms zero-shot state-of-the-art large language models by 20% in mIoU.