Learnability-Driven Submodular Optimization for Active Roadside 3D Detection

📅 2026-01-04
🏛️ arXiv.org
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work proposes LH3D, a learnability-oriented active learning framework for monocular 3D object detection from roadside cameras, addressing the challenges of annotation difficulty and low model learnability. LH3D introduces learnability as the core sampling criterion, explicitly suppressing inherently ambiguous and hard-to-annotate samples. By integrating submodular optimization, the framework enhances both annotation efficiency and model performance while ensuring informative coverage. Evaluated on the DAIR-V2X-I dataset, LH3D achieves 86.06% (vehicles), 67.32% (pedestrians), and 78.67% (cyclists) of the fully supervised performance using only 25% of the annotation budget, significantly outperforming conventional uncertainty-based active learning approaches.

Technology Category

Machine Learning: Active LearningComputer Vision: Learning & Optimization for CVSearch and Optimization: Learning to Search

Application Category

Economics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labelingSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingGraph Algorithms and Modeling for the Web: Representation, reconstruction, and subgraph or motif discovery in Web-related graphs
📝 Abstract
Roadside perception datasets are typically constructed via cooperative labeling between synchronized vehicle and roadside frame pairs. However, real deployment often requires annotation of roadside-only data due to hardware and privacy constraints. Even human experts struggle to produce accurate labels without vehicle-side data (image, LIDAR), which not only increases annotation difficulty and cost, but also reveals a fundamental learnability problem: many roadside-only scenes contain distant, blurred, or occluded objects whose 3D properties are ambiguous from a single view and can only be reliably annotated by cross-checking paired vehicle--roadside frames. We refer to such cases as inherently ambiguous samples. To reduce wasted annotation effort on inherently ambiguous samples while still obtaining high-performing models, we turn to active learning. This work focuses on active learning for roadside monocular 3D object detection and proposes a learnability-driven framework that selects scenes which are both informative and reliably labelable, suppressing inherently ambiguous samples while ensuring coverage. Experiments demonstrate that our method, LH3D, achieves 86.06%, 67.32%, and 78.67% of full-performance for vehicles, pedestrians, and cyclists respectively, using only 25% of the annotation budget on DAIR-V2X-I, significantly outperforming uncertainty-based baselines. This confirms that learnability, not uncertainty, matters for roadside 3D perception.
Problem

Research questions and friction points this paper is trying to address.

learnability
roadside 3D detection
active learning
inherently ambiguous samples
monocular 3D object detection
Innovation

Methods, ideas, or system contributions that make the work stand out.

learnability-driven
active learning
roadside 3D detection
inherently ambiguous samples
submodular optimization
🔎 Similar Papers
No similar papers found.