Object Detection Benchmarks are Incomplete: The Role of Label Errors and Annotation Uncertainty
论文指出标注不完整限制了目标检测基准质量,通过重新标注增加对象数量,并提出高召回率和不确定性感知的标注流程以提高数据质量。
论文指出标注不完整限制了目标检测基准质量,通过重新标注增加对象数量,并提出高召回率和不确定性感知的标注流程以提高数据质量。
BLInD通过车辆状态历史预测未来轨迹分布,采用AR和流匹配方法,减少AEB误报,实现低延迟实时部署。
This work addresses the challenge of sparse rewards and long-horizon exploration in 3D environments, where existing curiosity-driven methods often fall into local loops or redundant exploration. The authors propose a novel approach that integrates an online 3D reconstruction-based world model with RGB sequence-based policy learning, introducing spatial persistence and episodic context into the curiosity mechanism for the first time. This enables the agent to distinguish genuinely novel regions and avoid revisiting uninformative areas. Relying solely on visual inputs, the method achieves zero-shot generalization from Habitat-Matterport 3D (HM3D) to both Gibson and AI-generated environments after pure curiosity-driven pre-training, significantly outperforming from-scratch baselines and demonstrating strong performance on downstream tasks such as apple picking and image-goal navigation.
This work addresses the heavy reliance of camera pose estimation on large-scale 3D-annotated data by proposing a self-supervised pretraining approach that, for the first time, integrates inverse and forward dynamics models into this task. By learning latent action representations from large-scale unlabeled driving videos and using them as input features to a fine-tuned pose estimator, the method substantially reduces dependence on annotated data. With only a small amount of high-quality 3D annotations, it achieves over a 10% improvement in pose accuracy compared to state-of-the-art feedforward methods on both the Waymo and PandaSet benchmarks, attaining leading performance with an order of magnitude fewer annotations.
Existing 3D scene reconstruction methods struggle to meet the demands of interactive graphics applications for high-quality, cross-view consistent semantic segmentation. This work proposes a novel approach based on Radiant Foam—a voxelized Voronoi grid—by introducing an explicit semantic feature field at the cell level and, for the first time, directly incorporating spatial regularization into this field to achieve joint spatial-semantic decomposition. This design significantly enhances cross-view consistency and effectively mitigates artifacts caused by occlusions and inconsistent supervision. Experimental results demonstrate that the proposed method outperforms state-of-the-art approaches such as Gaussian Grouping and SAGA in both object-level semantic segmentation accuracy and consistency.
论文指出标注不完整限制了目标检测基准质量,通过重新标注增加对象数量,并提出高召回率和不确定性感知的标注流程以提高数据质量。
BLInD通过车辆状态历史预测未来轨迹分布,采用AR和流匹配方法,减少AEB误报,实现低延迟实时部署。
This work addresses the challenge of sparse rewards and long-horizon exploration in 3D environments, where existing curiosity-driven methods often fall into local loops or redundant exploration. The authors propose a novel approach that integrates an online 3D reconstruction-based world model with RGB sequence-based policy learning, introducing spatial persistence and episodic context into the curiosity mechanism for the first time. This enables the agent to distinguish genuinely novel regions and avoid revisiting uninformative areas. Relying solely on visual inputs, the method achieves zero-shot generalization from Habitat-Matterport 3D (HM3D) to both Gibson and AI-generated environments after pure curiosity-driven pre-training, significantly outperforming from-scratch baselines and demonstrating strong performance on downstream tasks such as apple picking and image-goal navigation.
This work addresses the heavy reliance of camera pose estimation on large-scale 3D-annotated data by proposing a self-supervised pretraining approach that, for the first time, integrates inverse and forward dynamics models into this task. By learning latent action representations from large-scale unlabeled driving videos and using them as input features to a fine-tuned pose estimator, the method substantially reduces dependence on annotated data. With only a small amount of high-quality 3D annotations, it achieves over a 10% improvement in pose accuracy compared to state-of-the-art feedforward methods on both the Waymo and PandaSet benchmarks, attaining leading performance with an order of magnitude fewer annotations.
Existing 3D scene reconstruction methods struggle to meet the demands of interactive graphics applications for high-quality, cross-view consistent semantic segmentation. This work proposes a novel approach based on Radiant Foam—a voxelized Voronoi grid—by introducing an explicit semantic feature field at the cell level and, for the first time, directly incorporating spatial regularization into this field to achieve joint spatial-semantic decomposition. This design significantly enhances cross-view consistency and effectively mitigates artifacts caused by occlusions and inconsistent supervision. Experimental results demonstrate that the proposed method outperforms state-of-the-art approaches such as Gaussian Grouping and SAGA in both object-level semantic segmentation accuracy and consistency.