Score
Designs and executes calibration audits and correction models that quantify systematic bias and offsets in reference datasets—especially computed references such as DFT outputs—relative to experimental benchmarks. Builds procedures to estimate reference offsets, evaluate uncertainty and validity of additive (or alternative) calibration schemes, and pre-register or reproduce calibration audits against benchmark sets.
This paper addresses the reliability of calibration evaluation for machine learning models, identifying systematic biases in the widely used Expected Calibration Error (ECE) under distributional shift and varying binning strategies. Methodologically, it clarifies the logical hierarchy among multi-level calibration definitions, and systematically exposes ECE’s limitations through visualization, binning-based statistical analysis, and theoretical derivation—demonstrating its failure to satisfy key requirements of robustness and consistency in calibration assessment. Building on this critique, the paper introduces and explicates emerging calibration paradigms—including distribution-level and instance-level calibration—alongside their corresponding evaluation methodologies, thereby constructing a rigorous, interpretable, and practice-oriented calibration knowledge framework. The results equip researchers with principled guidance for selecting appropriate evaluation metrics and advance calibration assessment from ad hoc, heuristic practices toward standardization and formalization.
This study addresses a critical gap in quantum software research: the absence of a systematic auditing mechanism for empirically grounded comparative claims, which has led to a pervasive “instantiation gap” characterized by insufficient evidentiary support. To bridge this gap, the authors propose CLAIMSTAB-QC, the first source-bound auditing framework tailored to empirical comparisons in quantum software. By integrating claim modeling, audit scope delimitation, evidence boundary identification, and directional classification, the framework enables precise validation of comparative assertions against original source materials. An evaluation across 455 claims from 119 papers reveals that only eight claims possessed sufficient matched evidence for direct auditing; among these, two were confirmed, four lacked adequate support, and two were contradicted—highlighting substantial deficiencies in the empirical rigor of current quantum software studies.
Existing multicalibration metrics lack simultaneous auditability and interpretable quantification of required model modifications. Method: We propose two novel information-theoretic, audit-friendly distance measures—differentiable Multicalibration (dMC) and Intersection Multicalibration—both rigorously equivalent to the geometric distance from the predictor to the set of perfectly multicalibrated predictors. These measures precisely quantify the magnitude of necessary adjustments and support verifiable auditing within the differentiable Calibration Error (dCE) framework. Contribution/Results: Theoretical analysis characterizes the loss landscape structure of multicalibration error and the geometric properties of its optimal solution set. Unlike prior generalizations, our metrics are the first to jointly satisfy auditability, reasonableness, and interpretability of modification magnitude. They establish a rigorous foundation for assessing multi-group fairness and robust uncertainty modeling, offering both theoretical insight and practical utility.
This study addresses the fundamental ambiguity in existing contamination audits, which often conflate truly clean data with insufficient detection power. To resolve this, the authors propose a theoretical framework based on sparse mixture models of the form \( Q_\alpha = (1-\alpha)P_0 + \alpha P_1 \), leveraging control-group estimation to assess statistical power. By integrating sample-splitting certificates with a two-stage planner, the method enables calibrated and reliable auditing. Key contributions include establishing distribution-free lower-bound certificates for contamination proportion, uncovering the failure mechanism of Gaussian budget calibration in small-sample regimes, and providing a corrective solution. Empirical results demonstrate high predictive accuracy of power curves (\( R^2 = 0.83\text{–}0.98 \)) across six channels and successfully reproduce the sensitivity ranking of injected contaminations: verbatim > paraphrase > surface. The repaired budgeting scheme proves conservatively effective, and non-rejection conclusions require joint evaluation of power, budget, and validity gating.
This study addresses the frequent failure of machine learning models in screening cathode materials for sodium-ion batteries, which often stems from reliance on computationally derived reference voltages—such as those from PBE+U—that exhibit systematic errors. Employing a preregistered experimental validation protocol, we evaluate a graph neural network model against experimental measurements across six known materials and quantify the discrepancy between computed and empirical voltages. Our results reveal that Materials Project’s PBE+U voltages are systematically underestimated by approximately 0.54 V, constituting the dominant source of model error (with a holdout-set MAE of 0.67 V and a 95% confidence upper bound of 1.09 V), and this bias strongly correlates with the voltage magnitude. To mitigate such issues, we propose a calibration and auditing framework tailored to DFT-based data ledgers, establishing a more reliable benchmark for computational materials screening.
This study addresses the limited reliability of existing training data contamination detection methods in real-world auditing scenarios, particularly when distribution shifts occur or when reference benchmarks are substantially smaller than the pretraining corpus. Through a systematic evaluation of three dominant paradigms—LLM Dataset Inference, Post-Hoc Dataset Inference, and CoDeC—the authors conduct 335 experiments across 27 open-source and state-of-the-art closed-source language models (up to 27B parameters). They identify distribution shift and small-scale benchmarks as two critical failure modes, revealing that only 199 evaluations yield correct conclusions. Current approaches suffer from high false-positive rates, low statistical power, or coarse-grained provenance resolution, rendering them inadequate for reliably verifying individual benchmark subsets and underscoring the irreplaceable value of transparent data provenance.
This work addresses the high cost of ground-truth evaluation in chemical and materials design, where existing machine learning surrogate models often lack reliability guarantees. Departing from conventional reliance on prediction accuracy metrics such as R²—which can paradoxically increase the risk of worst-case selections—the study proposes “rank preservation” as a core criterion for surrogate validation. It formally introduces the concept of “selection tax” and derives its theoretical upper and lower bounds. A safety certification framework for surrogates is established through selection-aware auditing, rank correlation analysis, and multi-task ground-truth validation. Experiments demonstrate that the proposed audit statistics achieve Spearman correlations of 0.80–0.99 with actual search performance, substantially outperforming R² (as low as 0.33). Certified screening strategies based on this framework reduce evaluation costs by up to 25-fold.
This work addresses the challenge that probabilistic programs generated by language models often suffer from statistical misspecifications—such as incorrect likelihoods, priors, or parameterizations—that are difficult to detect with conventional unit tests. The paper introduces, for the first time, Bayesian calibration as a central criterion for assessing the correctness of probabilistic programs and proposes a fully unsupervised, reference-free framework for their detection and repair. By integrating Bayesian validation techniques—including posterior predictive checks, simulation-based calibration (SBC), sampling diagnostics (e.g., $\hat{R}$, divergences, effective sample size), and held-out predictive log density—the method generates feedback signals to drive an iterative repair loop within large language models. Evaluated on 200 instances, the approach achieves detection AUCs of 0.97 with reference programs and 62–78% without, substantially outperforming unit testing; repair success rates reach 92% and 100% using GPT-5.1 and Claude, respectively.