Combining Foundation Model Confidence and Monocular Depth for Training-Free Out-of-Distribution Segmentation
提出了一种无需训练的方法,利用基础模型的置信度预测和单目深度估计来解决未知物体的检测与分割问题。
提出了一种无需训练的方法,利用基础模型的置信度预测和单目深度估计来解决未知物体的检测与分割问题。
研究提出CounterPlay方法,通过反事实自我对弈回溯失败任务并尝试不同驾驶风格以提高自动驾驶策略性能。
本文提出Diffuse2Seg方法,利用文本到图像扩散模型自动生成多粒度实例掩码,无需监督即可实现开放世界实体分割。
为解决端到端驾驶系统在策略引发状态下的安全问题,本文提出RoG-DAgger方法,通过短时运动预测构建高质量专家示例并适时调整控制权,以改善模型性能。
Current autonomous driving evaluation methods rely on aggregate metrics that struggle to effectively capture system failures in high-risk, low-frequency scenarios. To address this limitation, this work proposes RISC—a risk-informed evaluation protocol that enables model-agnostic and interpretable stress testing through computable risk slices, lightweight data annotation, and risk-guided sampling. Furthermore, the framework leverages large language models to assist in identifying critical yet often overlooked scenarios. Evaluated on monocular pedestrian perception tasks, RISC dramatically improves the detection rate of critical failures from 34.0% to 98.5%, demonstrating its superior capability in efficiently uncovering high-risk system deficiencies.
提出了一种无需训练的方法,利用基础模型的置信度预测和单目深度估计来解决未知物体的检测与分割问题。
研究提出CounterPlay方法,通过反事实自我对弈回溯失败任务并尝试不同驾驶风格以提高自动驾驶策略性能。
本文提出Diffuse2Seg方法,利用文本到图像扩散模型自动生成多粒度实例掩码,无需监督即可实现开放世界实体分割。
为解决端到端驾驶系统在策略引发状态下的安全问题,本文提出RoG-DAgger方法,通过短时运动预测构建高质量专家示例并适时调整控制权,以改善模型性能。
Current autonomous driving evaluation methods rely on aggregate metrics that struggle to effectively capture system failures in high-risk, low-frequency scenarios. To address this limitation, this work proposes RISC—a risk-informed evaluation protocol that enables model-agnostic and interpretable stress testing through computable risk slices, lightweight data annotation, and risk-guided sampling. Furthermore, the framework leverages large language models to assist in identifying critical yet often overlooked scenarios. Evaluated on monocular pedestrian perception tasks, RISC dramatically improves the detection rate of critical failures from 34.0% to 98.5%, demonstrating its superior capability in efficiently uncovering high-risk system deficiencies.