Offboard Occupancy Refinement with Hybrid Propagation for Autonomous Driving

📅 2024-03-13
📈 Citations: 2
✨ Influential: 1
📄 PDF
🤖 AI Summary
To address key challenges in vision-based 3D semantic scene completion (SSC)—including inaccurate joint geometric-semantic estimation, severe single-view occlusion, and inter-view discontinuities—this paper proposes OccFiner. Our method employs offline multi-view fusion, integrating a hybrid propagation mechanism comprising many-to-many local propagation and region-center global propagation, augmented by multi-view implicit alignment, explicit geometric modeling, and sensor bias correction. We introduce the first fully visual, automatic SSC annotation paradigm, enabling the first closed-loop visual occupancy data pipeline and city-scale static map refinement. On SemanticKITTI, OccFiner achieves state-of-the-art performance: its visual-only model matches LiDAR-based methods in accuracy, with significantly improved voxel-wise precision at long ranges. Comprehensive evaluation demonstrates the superiority of our offline multi-view approach across both geometric and semantic metrics.

Technology Category

Computer Vision: Multi-modal VisionIntelligent Robots: Multimodal Perception & Sensor FusionMachine Learning: Multi-instance/Multi-view Learning

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsEconomics, Online Markets and Human Computation: LLM based quality controls for crowd workGraph Algorithms and Modeling for the Web: Algorithms and analysis for incomplete, noisy, or partially observed Web-related graphs
📝 Abstract
Vision-based occupancy prediction, also known as 3D Semantic Scene Completion (SSC), presents a significant challenge in computer vision. Previous methods, confined to onboard processing, struggle with simultaneous geometric and semantic estimation, continuity across varying viewpoints, and single-view occlusion. Our paper introduces OccFiner, a novel offboard framework designed to enhance the accuracy of vision-based occupancy predictions. OccFiner operates in two hybrid phases: 1) a multi-to-multi local propagation network that implicitly aligns and processes multiple local frames for correcting onboard model errors and consistently enhancing occupancy accuracy across all distances. 2) the region-centric global propagation, focuses on refining labels using explicit multi-view geometry and integrating sensor bias, particularly for increasing the accuracy of distant occupied voxels. Extensive experiments demonstrate that OccFiner improves both geometric and semantic accuracy across various types of coarse occupancy, setting a new state-of-the-art performance on the SemanticKITTI dataset. Notably, OccFiner significantly boosts the performance of vision-based SSC models, achieving accuracy levels competitive with established LiDAR-based onboard SSC methods. Furthermore, OccFiner is the first to achieve automatic annotation of SSC in a purely vision-based approach. Quantitative experiments prove that OccFiner successfully facilitates occupancy data loop-closure in autonomous driving. Additionally, we quantitatively and qualitatively validate the superiority of the offboard approach on city-level SSC static maps. The source code will be made publicly available at https://github.com/MasterHow/OccFiner.
Problem

Research questions and friction points this paper is trying to address.

Enhancing vision-based 3D occupancy prediction accuracy
Addressing geometric and semantic estimation challenges
Improving multi-view continuity and occlusion handling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Offboard framework enhances vision-based occupancy accuracy
Hybrid propagation: multi-to-multi local and region-centric global
Achieves automatic SSC annotation in vision-only approach
🔎 Similar Papers
No similar papers found.