Score
Designs and implements procedures, instruments, or protocols to quantify variables and collect empirical data, specifying how measurements are taken and recorded. Builds and analyzes calibration, error propagation, uncertainty estimation, validation, and data‑quality processes to ensure measurements are accurate, precise, and reproducible.
This paper addresses the reliability of calibration evaluation for machine learning models, identifying systematic biases in the widely used Expected Calibration Error (ECE) under distributional shift and varying binning strategies. Methodologically, it clarifies the logical hierarchy among multi-level calibration definitions, and systematically exposes ECE’s limitations through visualization, binning-based statistical analysis, and theoretical derivation—demonstrating its failure to satisfy key requirements of robustness and consistency in calibration assessment. Building on this critique, the paper introduces and explicates emerging calibration paradigms—including distribution-level and instance-level calibration—alongside their corresponding evaluation methodologies, thereby constructing a rigorous, interpretable, and practice-oriented calibration knowledge framework. The results equip researchers with principled guidance for selecting appropriate evaluation metrics and advance calibration assessment from ad hoc, heuristic practices toward standardization and formalization.
In safety-critical applications, evaluating uncertainty calibration of regression models is hindered by inconsistent metric definitions, conflicting assumptions, and incomparable scales—impeding interpretability and reproducibility. This work systematically surveys and categorizes existing calibration metrics, then conducts a model-agnostic benchmark across real-world, synthetic, and manually miscalibrated datasets. We empirically demonstrate—for the first time—that most metrics yield contradictory or even opposing conclusions for identical calibration states, confirming that metric choice critically influences research outcomes. To address this, we propose ENCE (Expected Normalized Calibration Error) and CWC (Weighted Coverage Confidence) as more robust and stable primary metrics. Experiments across diverse scenarios show that ENCE and CWC exhibit superior consistency, strong resilience to noise and distribution shifts, and enhanced interpretability. Our findings establish a reproducible methodological foundation for uncertainty calibration evaluation in regression.
Experimental reproducibility in Empirical Software Engineering (ESE) is hindered by a fundamental disconnect between idealized methodological assumptions—e.g., standardized protocols and controlled conditions—and researchers’ actual experimental practices. Method: We conducted a two-year ethnographic study involving participant observation, in-depth interviews, and content analysis of experimental artifacts across diverse ESE research teams. Contribution/Results: We identify four critical dimensions—activity diversity, role distribution, conceptual granularity, and domain perspective—in which real-world experimentation systematically deviates from textbook models. Based on these findings, we propose the first high-fidelity conceptual and process model grounded in empirical research practice, explicitly capturing the “practice gap” underlying irreproducibility. This model provides foundational evidence and design principles for developing next-generation reproducibility-support tools, methodological guidelines, and evaluation frameworks in ESE.
Data scientists frequently lack systematic guidance when operationalizing ambiguous concepts (e.g., “writing authenticity,” “medical need”) into model-ready proxy target variables. To address this, we conducted semi-structured interviews with 15 data scientists across education and healthcare domains, followed by cross-domain thematic coding. We propose the “assemblage metrics” framework, identifying five core design criteria: validity, simplicity, predictiveness, portability, and resource efficiency. Our analysis reveals an iterative, problem-reconstruction–driven practice in which target variables are dynamically negotiated through trade-offs among these criteria. This work offers the first systematic characterization of such trade-offs in proxy target construction. It contributes a theoretically grounded framework and methodological tools for HCI, CSCW, and machine learning communities to support principled, transparent, and trustworthy predictive modeling—bridging conceptual abstraction with operationalizable measurement.
Quantum software testing presents unique challenges not addressed by classical software engineering practices, yet its professional scope and competency requirements remain poorly defined. Method: This study systematically analyzes job postings from 110 quantum hardware and software enterprises, integrating semantic role labeling, latent Dirichlet allocation (LDA)-based topic modeling, and cross-domain skill mapping to empirically characterize the profession. Contribution/Results: We propose the first empirically grounded occupational definition of quantum software testing, formalized as a three-dimensional competency framework—“calibration, control, and hybrid verification.” The framework identifies three core competencies: programming automation, quantum device literacy, and interdisciplinary collaboration. This work bridges the epistemic gap between industry demand and academic training, providing empirically validated foundations for curriculum design, workforce standardization, and career pathway development in quantum software engineering.
This study addresses the overreliance on inter-annotator agreement in current data annotation practices, which often overlooks annotation’s capacity to capture conceptual validity as a measurement act. Treating annotation as a measurement process, the work identifies five root causes of annotation issues—errors, ambiguity, impossibility, subjectivity, and annotator identity—and develops a measurement theory–based framework for diagnosing and improving annotation quality. Drawing on a synthesis of 132 literature sources and 10 semi-structured interviews, the research systematically defines target constructs, designs annotation instruments, implements labeling procedures, and evaluates both reliability and validity. The resulting framework equips annotation teams with evaluation methods that transcend mere agreement metrics, thereby substantially strengthening the foundational quality of AI training data.
This study addresses the challenge of disentangling sources of inter-laboratory variability—specifically baseline offsets versus differences in sensitivity—in multi-laboratory assessments of linear dose–response relationships. To this end, the authors propose a precision evaluation framework based on linear mixed-effects models, integrating analysis of variance, F-tests, and ISO 5725 standards to define and estimate repeatability and between-laboratory variance components. Overall measurement precision is quantified via average dose-specific variance. Under a fully balanced design, the framework yields an exact decomposition of total sum of squares and closed-form ANOVA estimators, overcoming the limitation of conventional fixed-effects models that detect only the presence of differences without identifying their origin. The approach was successfully applied to bronchoalveolar lavage fluid data from a rat intratracheal instillation study involving nanomaterials, effectively distinguishing the sources of observed variability.
Estimating the functional relationship between a continuous exposure and a binary outcome is challenging when covariates are measured with error. This study presents the first systematic evaluation of Simulation-Extrapolation, Regression Calibration, multiple imputation, and Bayesian correction methods, each coupled with flexible modeling techniques—including B-splines, P-splines, and fractional polynomials—within a multi-team, fully blinded, neutral simulation framework. By generating 155 distinct simulation scenarios and repeated samples, the research quantifies the bias and variance of each approach, revealing their relative strengths and limitations. The findings not only inform method selection under measurement error but also demonstrate the feasibility and value of this neutral comparative paradigm for rigorous methodological assessment.
This study addresses the limitations of traditional Gaussian process (GP) calibration methods, which neglect intermediate variables in computer experiments, leading to inadequate bias modeling and non-identifiability between the simulator and the discrepancy term. To resolve this, the authors propose a robust GP calibration framework that explicitly incorporates intermediate variables. The approach systematically selects key intermediate variables, constrains the discrepancy term using a scaled Gaussian stochastic process (S-GaSP), and employs space-filling designs to choose constraint points, thereby enabling identifiable joint modeling of the simulator and bias. This work is the first to systematically integrate intermediate variables into GP calibration, substantially improving predictive accuracy and the reliability of uncertainty quantification. Empirical results on nuclear binding energy prediction demonstrate clear superiority over existing baseline methods.