Stream3Dv2: Geometric-Semantic Fusion Enhanced Streaming Zero-Shot 3D Scene Understanding

📅 2026-08-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决零样本3D场景理解中流式RGB-D输入处理和噪声干扰问题,提出Stream3Dv2框架,通过几何-语义融合机制及点云优化策略实现高效准确的感知。
📝 Abstract
Recently, open-vocabulary zero-shot 3D scene understanding using vision foundation models has emerged as a promising alternative to data-intensive supervised methods. However, deploying these models in real-world scenarios is severely hindered by their inability to efficiently handle streaming RGB-D inputs and their inherent vulnerability to noise 2D segmentation masks. To address these critical limitations, we propose Stream3Dv2, a novel training-free framework designed for robust streaming 3D perception. Stream3Dv2 processes sequential data through an original nested local-to-historical architecture, capturing multi-view consistency while circumventing the high computational overhead so as to support timely responses. At its core, we introduce a comprehensive geometric-semantic fusion mechanism that resolves geometric noise and semantic ambiguity by explicitly utilizing semantic guidance and formulating 3D segmentation as solving point-and-set merging and partitioning problems. Furthermore, we present an innovative manifold-distance-based point cloud refinement strategy. This approach leverages local manifold graphs for point-to-manifold optimization that mitigates the boundary delineation failures caused by Euclidean-distance metrics, and employs geometric bounding boxes to dynamically activate and update historical instances for achieving rapid manifold-to-manifold refinement. Extensive experiments on public datasets demonstrate that Stream3Dv2 consistently outperforms existing baselines in foundational open-vocabulary streaming 3D segmentation and detection. Finally, we show that integrating our framework with an LLM-based agent enables advanced language-driven 3D scene understanding, underscoring its potential for open-world embodied intelligence. Code will be updated at https://github.com/SubmissionsIn/Stream3D.
Problem

Research questions and friction points this paper is trying to address.

zero-shot 3D scene understanding
streaming RGB-D inputs
noise 2D segmentation masks
Innovation

Methods, ideas, or system contributions that make the work stand out.

geometric-semantic fusion
streaming 3D perception
manifold-distance-based refinement
open-vocabulary zero-shot 3D scene understanding
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jie Xu
ISTD Pillar, Singapore University of Technology and Design, Singapore 487372
Na Zhao
Na Zhao
Singapore University of Technology and Design
Computer VisionMachine LearningScene Understanding3D PerceptionMultimedia