🤖 AI Summary
To address key challenges in vision-based 3D semantic scene completion (SSC)—including inaccurate joint geometric-semantic estimation, severe single-view occlusion, and inter-view discontinuities—this paper proposes OccFiner. Our method employs offline multi-view fusion, integrating a hybrid propagation mechanism comprising many-to-many local propagation and region-center global propagation, augmented by multi-view implicit alignment, explicit geometric modeling, and sensor bias correction. We introduce the first fully visual, automatic SSC annotation paradigm, enabling the first closed-loop visual occupancy data pipeline and city-scale static map refinement. On SemanticKITTI, OccFiner achieves state-of-the-art performance: its visual-only model matches LiDAR-based methods in accuracy, with significantly improved voxel-wise precision at long ranges. Comprehensive evaluation demonstrates the superiority of our offline multi-view approach across both geometric and semantic metrics.
📝 Abstract
Vision-based occupancy prediction, also known as 3D Semantic Scene Completion (SSC), presents a significant challenge in computer vision. Previous methods, confined to onboard processing, struggle with simultaneous geometric and semantic estimation, continuity across varying viewpoints, and single-view occlusion. Our paper introduces OccFiner, a novel offboard framework designed to enhance the accuracy of vision-based occupancy predictions. OccFiner operates in two hybrid phases: 1) a multi-to-multi local propagation network that implicitly aligns and processes multiple local frames for correcting onboard model errors and consistently enhancing occupancy accuracy across all distances. 2) the region-centric global propagation, focuses on refining labels using explicit multi-view geometry and integrating sensor bias, particularly for increasing the accuracy of distant occupied voxels. Extensive experiments demonstrate that OccFiner improves both geometric and semantic accuracy across various types of coarse occupancy, setting a new state-of-the-art performance on the SemanticKITTI dataset. Notably, OccFiner significantly boosts the performance of vision-based SSC models, achieving accuracy levels competitive with established LiDAR-based onboard SSC methods. Furthermore, OccFiner is the first to achieve automatic annotation of SSC in a purely vision-based approach. Quantitative experiments prove that OccFiner successfully facilitates occupancy data loop-closure in autonomous driving. Additionally, we quantitatively and qualitatively validate the superiority of the offboard approach on city-level SSC static maps. The source code will be made publicly available at https://github.com/MasterHow/OccFiner.