Score
Designs, implements, and evaluates procedures, protocols, and algorithms that align model outputs or sensor measurements with real-world values, including computing and comparing calibration metrics (e.g., expected calibration error, Brier score), implementing calibration methods (temperature scaling, classifier calibration), and assessing pooled, subgroup, decision- and difficulty-aware calibration. Also develops and applies calibration set and protocol designs and performs hardware sensor calibration and alignment tasks such as camera intrinsics/extrinsics and camera–lidar calibration, and compares calibration across models or modalities.
Probabilistic outputs of AI models often exhibit miscalibration—i.e., predicted confidence scores poorly reflect true accuracy—hindering their reliable deployment in safety-critical applications and ensemble systems. Method: This paper presents a systematic survey of probabilistic calibration evaluation methods for classification and object detection models. Grounded in statistical assessment theory, it unifies diverse approaches—including reliability diagrams, Brier score, expected calibration error (ECE), maximum calibration error (MCE), Kolmogorov–Smirnov test, ROC-based metrics, and IoU-aware measures—within a coherent framework covering binary, multiclass, and detection tasks. Contribution: We propose the first taxonomy of calibration metrics, categorizing 82 existing measures into four families: point-wise, binning-based, kernel/curve-based, and cumulative. Additionally, we introduce the first structured calibration metric knowledge base, enabling rapid metric selection, implementation, and comparative analysis—thereby establishing new interpretable and quantifiable benchmarks for trustworthy AI.
This paper addresses the reliability of calibration evaluation for machine learning models, identifying systematic biases in the widely used Expected Calibration Error (ECE) under distributional shift and varying binning strategies. Methodologically, it clarifies the logical hierarchy among multi-level calibration definitions, and systematically exposes ECE’s limitations through visualization, binning-based statistical analysis, and theoretical derivation—demonstrating its failure to satisfy key requirements of robustness and consistency in calibration assessment. Building on this critique, the paper introduces and explicates emerging calibration paradigms—including distribution-level and instance-level calibration—alongside their corresponding evaluation methodologies, thereby constructing a rigorous, interpretable, and practice-oriented calibration knowledge framework. The results equip researchers with principled guidance for selecting appropriate evaluation metrics and advance calibration assessment from ad hoc, heuristic practices toward standardization and formalization.
Addressing the challenge of achieving both high accuracy and low computational overhead in camera–LiDAR extrinsic calibration for autonomous driving, this paper proposes a multi-objective optimization framework that explicitly models computational cost—including point cloud sampling rate and runtime resource consumption—as an optimization objective, jointly minimizing image-edge projection geometric error and embedded-deployment overhead. Leveraging NSGA-II, the framework simultaneously optimizes the 6-DoF pose parameters and sampling strategy, yielding an interpretable Pareto front on the KITTI dataset to enable preference-driven solution selection. Compared to gradient-based and learning-based methods, our approach achieves comparable calibration accuracy while significantly reducing inference latency and memory footprint. This provides an efficient, reliable, and tunable calibration solution tailored for resource-constrained platforms.
To address the poor robustness and deployment difficulty of LiDAR–camera extrinsic calibration in mass production and after-sales scenarios for autonomous driving, this paper proposes a fully automatic calibration method based on square calibration targets. The method introduces a purely geometry-driven, multi-stage target detection pipeline, a hierarchical coarse-search mechanism insensitive to initial pose errors, and a direct optimization algorithm incorporating photometric consistency constraints—collectively enhancing robustness against sensor noise, sparse or incomplete point clouds, and large initial misalignments. Crucially, it requires no specialized retroreflective materials and achieves rapid (<1 minute), high-precision calibration (rotational error <0.1°, translational error <2 mm). Extensive experiments validate its stability and deployability in real-world manufacturing lines and after-sales service environments, demonstrating strong scalability for large-scale deployment of multi-sensor systems in production settings.
This work addresses the boundary blurring and distortion in LiDAR–camera extrinsic calibration caused by laser beam footprint effects and mixed-intensity returns. To this end, the authors propose a joint calibration method that integrates boundary response modeling. By co-observing visual fiducials on a printable planar calibration board and LiDAR-visible circular reflective boundaries, the approach iteratively refines 3D LiDAR edge feature points. It further incorporates an intensity- and geometry-constrained refinement strategy and a confidence-weighted reprojection optimization framework. Notably, this is the first method to embed explicit modeling of LiDAR boundary response characteristics into the extrinsic calibration pipeline. Evaluated on real-world data, it achieves sub-pixel reprojection accuracy and millimeter-level feature consistency, substantially improving downstream visual–LiDAR odometry performance.
Real-time detection of LiDAR–camera extrinsic parameter misalignment remains challenging in autonomous driving systems. Method: This paper proposes a lightweight end-to-end binary classification framework that reframes conventional parameter regression as a calibration-state discrimination task. We innovatively introduce contrastive learning into multimodal sensor calibration verification, designing a Siamese-network-based cross-modal feature embedding model with a lightweight CNN backbone and cosine similarity–based decision mechanism—eliminating reliance on geometric priors, specific object classes, or driving behavior. Contribution/Results: The method achieves >98% detection accuracy on both the KITTI benchmark and a custom dataset, with inference latency under 10 ms, enabling deployment on embedded platforms. It significantly outperforms state-of-the-art approaches and the source code is publicly available.
This work addresses the challenging problem of cross-modal calibration between LiDAR and cameras by proposing the first bird’s-eye-view (BEV)-based alignment framework. The method unifies multimodal data into a shared BEV space and employs a two-stage optimization strategy: it first implicitly regresses coarse calibration parameters and then explicitly aligns cross-modal features, enhanced by a CLIP-inspired contrastive loss to enforce semantic consistency. By integrating domain-specific BEV feature extraction with contrastive learning constraints, the approach significantly outperforms existing methods, achieving state-of-the-art calibration accuracy. On the KITTI and nuScenes benchmarks, it reduces relative rotation error by 51% and 68%, and translation error by 80% and 91%, respectively.
This work addresses the limitation of existing safety-critical systems, which typically evaluate only predictive accuracy while lacking rigorous validation of the overall calibration of predicted probability distributions. To bridge this gap, the authors propose a modular calibration testing framework that decouples the calibration process into four interchangeable components: data model, scoring rule, hypothesis formulation, and statistical test procedure. Built upon formal statistical hypothesis testing, the framework provides a single accept/reject decision for the entire predictive distribution. Crucially, it rejects only overly confident predictions while tolerating reasonable deviations, thereby balancing practicality with flexibility. Empirical evaluations on weather forecasting and robotic pose estimation tasks demonstrate that the framework effectively supports reliable deployment in safety-critical applications.
This study addresses the limitation of existing posterior calibration evaluation practices, which predominantly focus on aggregate performance while overlooking robustness across varying operational conditions within datasets. The authors propose the first conditionally stratified evaluation framework, employing preregistered experiments to systematically compare the robustness of temperature scaling (TEMP) and isotonic regression (ISO) under four controlled conditions. The assessment encompasses multiple dimensions—including calibration slope, Brier score, discrimination, and AUROC—and applies Holm’s correction for multiplicity in rigorous hypothesis testing. Results demonstrate that TEMP consistently achieves calibration slopes closer to one and superior, more stable Brier scores across all conditions; differences in discrimination between the two methods are negligible; and AUROC performance varies by condition, revealing that the robustness of calibration methods is highly dependent on both specific operational contexts and the choice of evaluation metric.
This work proposes a ray-based camera calibration framework tailored for 3D reconstruction, addressing the limitations of traditional reprojection error–based methods that rely on 2D calibration boards and inadequately reflect 3D geometric accuracy. Instead of reprojection error, the approach introduces reconstruction error and intersection error as more representative metrics. It employs a novel icosahedral 3D calibration target and a ring-shaped feature detector, integrated with a generalized distortion model and bootstrapping to refine both intrinsic and extrinsic parameter estimates. Experimental results on synthetic data demonstrate that the proposed method reduces average intersection error by approximately 40%, significantly enhances calibration stability, and validates that ray-level metrics provide a more faithful assessment of 3D reconstruction fidelity compared to conventional approaches.