Score
Designs and implements methods that convert object-detection performance metrics (precision, recall, confidence, detection range) into quantitative estimates of scene visibility (for example visibility distance, transmissivity, or a scalar visibility score) by establishing and validating mappings or proxies between detection outcomes and visibility measures. Builds and analyzes visibility estimators and uses detection-driven visibility outputs to evaluate or compare sensing and image-enhancement pipelines by measuring how changes in detection performance affect downstream tasks such as navigation or decision-making.
This study addresses the lack of physically interpretable linkage between existing image dehazing quality metrics and actual visibility distance, which hinders their utility in maritime navigation safety decisions. To bridge this gap, the authors construct a Maritime Simulated Visibility Dataset (MSVD) using Unity3D and propose a visibility-oriented evaluation framework that leverages object detection accuracy as an intermediary proxy. This framework uniquely maps dehazing performance directly to quantifiable gains in visual range, establishing a novel visibility metric with both physical interpretability and operational safety relevance. The approach is validated across diverse imaging conditions, demonstrating consistent reliability. Notably, MSVD serves as the first maritime simulation benchmark with precisely annotated visibility distances, significantly enhancing the quantitative assessment of dehazing algorithms in terms of navigational safety and operational efficiency.
Existing segmentation evaluation metrics often lack transparency and modularity, making them ill-suited for diverse tasks such as transparent object, specular surface, or lesion segmentation. This work proposes a unified evaluation framework that decomposes metrics into five modular components: prediction representation, target extraction, target matching, score computation, and metric reporting. For the first time, it systematically analyzes the implicit assumptions and design limitations of mainstream binary segmentation metrics through this modular lens. The framework enables task-aware customization of evaluation protocols, reveals evolutionary trajectories among existing metrics, and is accompanied by an open-source toolkit. By offering a principled and interpretable foundation, this approach paves the way for developing more rational and adaptable segmentation evaluation methodologies.
This work addresses the insufficient robustness of autonomous driving perception systems under adverse or adversarial driving conditions, which often leads to detection errors or latency. The authors propose a novel framework for quantifying perceptual sensitivity by integrating model disagreement and inference variability, leveraging an ensemble of five state-of-the-art vision models—YOLOv8, YOLOv9, DETR50, DETR101, and RT-DETR—to systematically evaluate performance degradation in both simulated and real-world scenarios involving low illumination, fog, occlusion, and long-range detection. Innovatively, they introduce a stopping-distance-based perceptual evaluation metric that reveals the compounded negative impact of multiple concurrent factors on perception robustness, thereby establishing a new paradigm for safety validation in autonomous driving systems.
This study investigates whether vision-language models can accurately assess the visibility of image content and proactively abstain when evidence is insufficient. To this end, the authors construct a benchmark comprising 300 rigorously curated evaluation items, leveraging minimally edited image–text pairs to probe models’ judgment and abstention capabilities. They introduce “reason encoding” to explain unanswerability and employ controlled minimal edits to verify whether model responses shift appropriately with changes in evidential support. The work proposes a suite of multidimensional evaluation metrics: Confidence-Aware Accuracy (CAA), Minimal-Edit Flip Rate (MEFR), Selective Prediction via Confidence Ranking (SelRank), and Theory-of-Mind-inspired Accuracy (ToMAcc). Experiments reveal that GPT-4o and Gemini 1.5 Pro achieve the best overall performance (aggregate score ≈ 0.728), while the open-source Gemma 3 12B model (0.505) outperforms several older closed-source systems.
Deep image quality assessment (IQA) models lack systematic evaluation of perceptual invariance to affine transformations—such as rotation, translation, scaling, and spectral illuminant changes—despite the human visual system’s robustness to such geometric and photometric variations. Method: We introduce a psychophysics-based quantification framework grounded in two-alternative forced-choice experiments to define and measure “imperceptibility thresholds” for affine distortions in a unified metric space, enabling direct comparison between model outputs and human perception. Contribution/Results: This is the first work to systematically calibrate and benchmark affine invariance thresholds across leading deep IQA models (e.g., LPIPS, DISTS). Our framework is model-agnostic and generalizable. Experiments reveal that all state-of-the-art deep IQA models exhibit significant deviations from human-level invariance behavior, indicating that optimizing solely for distortion visibility fails to capture the human visual system’s essential structural invariance mechanisms. The study establishes a new perceptual alignment benchmark and an interpretable calibration paradigm for IQA model evaluation.
Existing perception evaluation metrics (e.g., Precision/Recall/F1) emphasize aggregate accuracy while neglecting the safety-critical impact of false positives—particularly hazardous misclassifications that may trigger severe autonomous driving accidents, even in high-scoring systems. To address this fundamental gap, we propose EPSM—the first safety-oriented environmental perception metric that jointly models risk coupling between object detection and lane-line detection. EPSM introduces lightweight object- and lane-safety measures, incorporates risk-sensitive false-positive cost modeling, and performs cross-task safety dependency analysis. Empirical evaluation on the DeepAccident dataset demonstrates that EPSM effectively identifies numerous “high-accuracy, high-risk” false detections, significantly improving detection of catastrophe-inducing errors. It thereby bridges a critical blind spot of conventional accuracy metrics in accident prevention, providing an interpretable, quantifiable theoretical framework and practical tool for safety-driven perception system evaluation.
This work addresses the challenge that existing large vision-language models struggle to accurately perceive and describe low-level physical degradations in remote sensing images due to domain shift. To bridge this gap, the authors introduce SenseBench, the first benchmark for low-level visual diagnosis in remote sensing, grounded in a physics-driven hierarchical taxonomy encompassing six major categories and 22 fine-grained degradation types, with over 10,000 meticulously annotated samples. The benchmark features a dual-task evaluation protocol assessing both perception and description capabilities. Systematic evaluation of 29 state-of-the-art vision-language models reveals critical limitations, including domain shift, confusion among multiple concurrent degradations, fluent but hallucinated descriptions, and inversion between perception and description. SenseBench thus provides a high-quality dataset and a reliable evaluation platform to advance research on remote sensing image quality understanding.
This study addresses the widespread misuse of confidence scores from open-vocabulary detectors—such as Grounding DINO, OWLv2, and SAM3—as proxies for object visibility, when in fact these scores reflect category presence rather than the actual visibility of a specific instance. By constructing a ground-truth visibility benchmark using a geometric segmentation oracle and evaluating across multiple simulated environments and real-world video data, the work systematically audits current models and reveals, for the first time, that detector confidence remains high even when only 1/8 of an object is visible. This fundamental mismatch leads to a nearly tenfold underestimation of active perception performance. The authors argue that confidence-based evaluation and gating mechanisms are inherently biased and advocate instead for object-anchored visibility signals. To support future research, they release the first controllable occlusion benchmark dataset.
This work addresses the high training cost associated with evaluating synthetic object detection datasets by introducing CCDM (Conditional-Composition Domain Match), the first family of training-free proxy metrics tailored for synthetic detection data. CCDM predicts the relative utility of synthetic data for downstream detectors by precomputing image-level and instance-level conditional composition and domain-matching similarities. Evaluated on VisDrone-DET, CCDM achieves a Spearman correlation coefficient of 1.0 with YOLOv8 performance, substantially outperforming existing evaluation methods and significantly enhancing the efficiency of synthetic data selection.
Existing object detection models lack intuitive, fine-grained methods for performance comparison, making it difficult to uncover their shared and distinct failure modes in recognizing ground-truth labels. To address this, this work proposes Differences in Detection (DnD), a novel approach that introduces a structured set-partitioning mechanism based on standard matching algorithms. By decomposing model behaviors into intersections, differences, and co-missed sets, DnD enables direct pairwise comparison and integrates the TIDE error taxonomy to construct an interpretable confusion matrix. Moving beyond conventional metrics like mAP and isolated error statistics, the method clearly delineates shared versus unique errors, thereby guiding interpretability techniques—such as ODAM—to prioritize critical samples that reveal meaningful discrepancies between detectors.