measurement-simulation alignment

Designs and implements methods and pipelines to compare, correlate, and align experimental or laboratory measurements with simulation outputs; builds statistical and algorithmic tools such as calibration functions, mapping/transformation models, uncertainty and error models, and agreement metrics to quantify and reconcile differences. Analyzes paired measurement–simulation datasets to detect biases and systematic errors, validate models, estimate uncertainty, and produce calibrated simulation parameters or adjusted measurement procedures.

measurement-simulationalignment

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$190K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This paper addresses the reliability of calibration evaluation for machine learning models, identifying systematic biases in the widely used Expected Calibration Error (ECE) under distributional shift and varying binning strategies. Methodologically, it clarifies the logical hierarchy among multi-level calibration definitions, and systematically exposes ECE’s limitations through visualization, binning-based statistical analysis, and theoretical derivation—demonstrating its failure to satisfy key requirements of robustness and consistency in calibration assessment. Building on this critique, the paper introduces and explicates emerging calibration paradigms—including distribution-level and instance-level calibration—alongside their corresponding evaluation methodologies, thereby constructing a rigorous, interpretable, and practice-oriented calibration knowledge framework. The results equip researchers with principled guidance for selecting appropriate evaluation metrics and advance calibration assessment from ad hoc, heuristic practices toward standardization and formalization.

Evaluation MetricsLimitationsMachine Learning Calibration

Simulations in Statistical Workflows

Mar 31, 2025
PB
Paul-Christian Burkner
🏛️ TU Dortmund University | Independent Scientist | Rensselaer Polytechnic Institute

This paper systematically examines the structural role and evolutionary trajectory of simulation methods across the statistical lifecycle. Addressing the current fragmentation and conceptual ambiguity in simulation practice, the study introduces, for the first time, a comprehensive functional taxonomy—spanning model specification, diagnostic checking, validation, and inference—and proposes a “simulation-driven” paradigm for statistical practice, prioritizing computational scalability. Methodologically, it integrates Monte Carlo simulation, approximate Bayesian computation (ABC), simulation-based calibration, and posterior predictive checking, implemented via high-performance computing frameworks to enable large-scale empirical analysis. Key contributions are: (1) establishing simulation as foundational statistical infrastructure; (2) providing an actionable roadmap for algorithm design, statistical software development, and pedagogical reform; and (3) advancing a paradigm shift in statistical practice—from model-centric to simulation-augmented inference.

Analyzing trends in simulation-based statistical algorithmsExamining simulation roles in statistical workflowsExploring future impacts of simulations on statistics

Evaluating the Quality of the Quantified Uncertainty for (Re)Calibration of Data-Driven Regression Models

Aug 25, 2025
JW
Jelke Wibbeke
🏛️ Carl von Ossietzky Universität Oldenburg | German Aerospace Center (DLR) | Jade University of Applied Science

In safety-critical applications, evaluating uncertainty calibration of regression models is hindered by inconsistent metric definitions, conflicting assumptions, and incomparable scales—impeding interpretability and reproducibility. This work systematically surveys and categorizes existing calibration metrics, then conducts a model-agnostic benchmark across real-world, synthetic, and manually miscalibrated datasets. We empirically demonstrate—for the first time—that most metrics yield contradictory or even opposing conclusions for identical calibration states, confirming that metric choice critically influences research outcomes. To address this, we propose ENCE (Expected Normalized Calibration Error) and CWC (Weighted Coverage Confidence) as more robust and stable primary metrics. Experiments across diverse scenarios show that ENCE and CWC exhibit superior consistency, strong resilience to noise and distribution shifts, and enhanced interpretability. Our findings establish a reproducible methodological foundation for uncertainty calibration evaluation in regression.

Evaluating conflicting calibration metrics for regression modelsIdentifying inconsistencies in recalibration metric performance comparisonsSystematically benchmarking reliability of uncertainty quantification methods

Estimating the functional relationship between a continuous exposure and a binary outcome is challenging when covariates are measured with error. This study presents the first systematic evaluation of Simulation-Extrapolation, Regression Calibration, multiple imputation, and Bayesian correction methods, each coupled with flexible modeling techniques—including B-splines, P-splines, and fractional polynomials—within a multi-team, fully blinded, neutral simulation framework. By generating 155 distinct simulation scenarios and repeated samples, the research quantifies the bias and variance of each approach, revealing their relative strengths and limitations. The findings not only inform method selection under measurement error but also demonstrate the feasibility and value of this neutral comparative paradigm for rigorous methodological assessment.

covariate adjustmentexposure-outcome relationshipfunctional form

This study addresses the unreliable estimation of repeatability, between-laboratory, and reproducibility variance components under ISO 5725 standards when sample sizes are small or variance structures are extreme. To overcome this limitation, the authors propose a tailored Bootstrap resampling strategy adapted to a one-way random effects model. The approach refines point estimates by adjusting within-laboratory resampling and constructs confidence intervals via a two-stage resampling scheme integrated with bias-corrected and accelerated (BCa) techniques. Extensive simulations and validation using real data from ISO 5725-4 demonstrate that the proposed method substantially improves estimation accuracy and confidence interval coverage. It yields reliable, near-nominal or conservatively valid inferences for small- to moderate-sized experiments and clearly delineates optimal strategies across different practical scenarios.

bootstrapinterlaboratory precisionISO 5725

Latest Papers

What's happening recently
View more

This study addresses the widespread lack of systematic training in instrumentation software and machine learning tools among early-career researchers in high-energy physics, a gap that significantly hinders their research efficiency and professional development. Focusing specifically on this cohort’s practical needs and deficiencies regarding open-source software and machine learning education, the project collected feedback from 174 early-career researchers through a structured survey and employed statistical analysis to evaluate the accessibility and quality of existing training programs. The findings reveal that approximately 70% of respondents have received no such training. These results provide empirical evidence to inform the design of targeted, effective training frameworks aimed at enhancing the computational and analytical competencies of young scientists in the field.

early-career researchersHEP instrumentationmachine learning

This study addresses the challenge of effectively validating input model specifications in digital twin simulations, where conventional approaches—relying solely on marginal output distributions—often fail to detect misspecified joint input models. To overcome this limitation, the authors propose a novel statistical validation framework based on sub-trajectory conditioning. By repeatedly restarting simulations from observed system states while conditioning on subsets of random inputs, the method constructs conditional output distributions that enable goodness-of-fit testing of the full joint input model. This approach innovatively transcends the constraints of marginal validation and is complemented by diagnostic tools to pinpoint specific input sources responsible for detected discrepancies. Empirical evaluations on M/M/1 and tandem queueing systems demonstrate the framework’s heightened sensitivity and effectiveness, successfully identifying input model misspecifications that traditional methods overlook.

conditional output distributiondigital twinsgoodness-of-fit

When machine learning models are employed as measurement instruments, it remains unclear whether their outputs genuinely reflect stable and consistent latent constructs beyond merely achieving predictive performance. This work formally introduces the concept of “learned measurement functions” and proposes “measurement stability” as a distinct evaluation criterion. Through theoretical analysis and empirical case studies, we demonstrate that conventional metrics—such as generalization error, calibration, and robustness—do not guarantee measurement consistency. Our findings reveal that models with comparable predictive accuracy can implement systematically inequivalent measurement functions, and that these discrepancies become pronounced under distributional shifts, thereby exposing critical limitations in current evaluation frameworks.

distribution shiftinductive biaslearned measurement

This work addresses the limitation of existing safety-critical systems, which typically evaluate only predictive accuracy while lacking rigorous validation of the overall calibration of predicted probability distributions. To bridge this gap, the authors propose a modular calibration testing framework that decouples the calibration process into four interchangeable components: data model, scoring rule, hypothesis formulation, and statistical test procedure. Built upon formal statistical hypothesis testing, the framework provides a single accept/reject decision for the entire predictive distribution. Crucially, it rejects only overly confident predictions while tolerating reasonable deviations, thereby balancing practicality with flexibility. Empirical evaluations on weather forecasting and robotic pose estimation tasks demonstrate that the framework effectively supports reliable deployment in safety-critical applications.

calibrationdistributional validationprobabilistic forecasting

This study addresses the challenge of disentangling sources of inter-laboratory variability—specifically baseline offsets versus differences in sensitivity—in multi-laboratory assessments of linear dose–response relationships. To this end, the authors propose a precision evaluation framework based on linear mixed-effects models, integrating analysis of variance, F-tests, and ISO 5725 standards to define and estimate repeatability and between-laboratory variance components. Overall measurement precision is quantified via average dose-specific variance. Under a fully balanced design, the framework yields an exact decomposition of total sum of squares and closed-form ANOVA estimators, overcoming the limitation of conventional fixed-effects models that detect only the presence of differences without identifying their origin. The approach was successfully applied to bronchoalveolar lavage fluid data from a rat intratracheal instillation study involving nanomaterials, effectively distinguishing the sources of observed variability.

between-laboratory variancedose-response relationshipinterlaboratory studies

Hot Scholars

JZ

Jeff Z. Pan

Professor of Knowledge Computing, University of Edinburgh
Artificial IntelligenceKnowledge Representation and ReasoningKnowledge Based Learning
HW

Haozhao Wang

Huazhong University of Science and Technology
Could-edge Distributed LearningFederated LearningAI SecurityMulti-modal LLM Agent
ZW

Zekun Wu

Research Scientist, Holistic AI / PhD Student, University College London
Agentic AIResponsible AIBehavioural RobustnessExplainability
PF

Paolo Favaro

Professor of Computer Vision, University of Bern
computer visionmachine learningcomputational photographyinverse problems
HJ

Hyoungwook Jin

University of Michigan
human-computer interactionend-learner programmingpersonalized education