Score
Designs and implements methods to measure, quantify, and interpret differences between predictions produced by multiple models, modalities, or model versions, including construction of discrepancy metrics and tests that summarize prediction gaps. Builds procedures that use these discrepancy measures to detect anomalous inputs or distribution shifts (e.g., derive OOD or risk scores, flag inputs with large prediction gaps) and evaluates and validates those signals on held-out data.
This work addresses the lack of systematic model misspecification diagnostics in simulation-based inference (SBI). We propose the first unified framework for multi-scale misspecification diagnosis. Methodologically, we develop a distortion-based statistical testing theory explicitly grounded in classical hypothesis testing; design a self-calibrating neural density estimation algorithm for end-to-end misspecification identification; and integrate distortion-driven testing, simulation-based Bayesian inference, and joint residual–outlier analysis. Our key contributions are: (i) the first interpretable and scalable testing paradigm bridging local anomaly detection to global model validation; and (ii) empirical validation across diverse simulation tasks and on the real gravitational-wave event GW150914—reproducing established results while successfully extending diagnostics to high-dimensional, complex forward models.
This study addresses structural conflicts—arising both among heterogeneous data sources and between data and model assumptions—in evidence synthesis models. We propose a general conflict detection framework based on score-based discrepancy measures. Methodologically, we extend prior–data conflict diagnostics to the latent space of hierarchical models, enabling inconsistency detection under multilevel and non-exchangeable structures; integrating Bayesian evidence synthesis, score-function-based metrics, and posterior simulation, our approach provides quantitative assessment of model assumption–data compatibility. Key contributions include: (1) moving beyond conventional bias diagnostics confined to the prior–likelihood level; (2) demonstrating high sensitivity to conflicts in both exchangeable and non-exchangeable models; and (3) exhibiting complementary diagnostic capability to existing methods in a real-world influenza severity model, thereby significantly enhancing the reliability of complex Bayesian inference.
Counterfactual prediction under evolving intervention policies or hypothetical decision scenarios remains challenging due to unobservable potential outcomes, hindering model identifiability, evaluation, and generalization. Method: We propose the first systematic theoretical framework addressing this challenge—comprising (i) identifiability conditions for counterfactual prediction models, (ii) a performance evaluation system targeting loss, AUC, and calibration, and (iii) robust hyperparameter selection under model misspecification. Our approach integrates causal inference principles, doubly robust estimation, and loss-driven evaluation metric design. Contribution/Results: Validated via simulation studies and a real-world clinical application—cardiovascular risk prediction in statin-naïve populations—the framework significantly improves out-of-distribution generalization and clinical decision reliability in counterfactual settings.
Existing methods struggle to align and interpret distribution shifts across heterogeneous, domain-consistent datasets—such as tabular, textual, visual, and time-series data—especially when scale and modality disparities are pronounced, resulting in poor interpretability. This paper introduces the first human-centric, cross-modal distribution discrepancy explanation framework, implemented as an interpretable dataset comparison toolbox. It integrates statistical hypothesis testing, feature importance decomposition, class activation mapping (CAM), contrastive representation learning, and interpretable generative modeling to enable fine-grained, semantically readable attribution and visualization of distributional shifts. Evaluated across diverse real-world scenarios, the framework significantly improves users’ efficiency in understanding shift causes and enhances the accuracy of intervention decisions—thereby overcoming the limitations of conventional black-box shift detection approaches.
Accurately evaluating the performance of analytical methods in simulation studies is hindered by underreporting and inconsistent handling of “missingness” issues—such as algorithm failure or non-convergence—that compromise validity and reproducibility. Method: We conducted a large-scale empirical analysis of 482 methodological simulation studies, systematically extracting metadata, applying qualitative coding, and performing case studies—including publication bias correction—to quantify the prevalence and reporting practices of missingness. Contribution/Results: We found that only 23% of studies mentioned missingness and merely 14% described mitigation strategies. Based on these findings, we developed a novel missingness taxonomy tailored to simulation research and proposed actionable principles—including mandatory missingness reporting—alongside a comprehensive, end-to-end practice guideline covering reporting, handling, and replication. Validation confirmed substantial improvements in transparency, comparability, and reproducibility of simulation studies.
Current evaluation practices for supervised learning models are often misleading due to an overreliance on single aggregate metrics, which neglect the alignment among data characteristics, task objectives, and real-world application contexts. This work reframes model evaluation as a context-dependent, decision-oriented process and systematically investigates—through controlled experiments—the impact of dataset properties, validation strategies, class imbalance, and asymmetric error costs on evaluation outcomes. Leveraging diverse benchmark datasets, multiple validation protocols, and multidimensional performance measures, the study uncovers common pitfalls such as the accuracy paradox, data leakage, and metric misuse. It proposes a structured evaluation framework explicitly aligned with operational goals, offering principled guidance for developing more robust, reliable, and trustworthy supervised learning systems.
This study addresses the problem of determining whether high-frequency monitoring data return to their pre-intervention baseline distribution following an intervention. The authors propose a sequential testing procedure that requires no assumptions about the underlying data distribution. The method constructs a discrepancy measure via universal inference and combines it with individualized empirical calibration to form a non-negative supermartingale, yielding an e-process that enables valid detection of the recovery time at any arbitrary stopping point without specifying a null model. Theoretical analysis provides finite-sample bounds on the calibration error, and both simulations and a clinical case study demonstrate the method’s superior performance in accurately identifying the time at which baseline conditions are restored.
This work addresses the practical challenge in industrial anomaly detection where “normal” samples often exhibit ambiguous definitions—such as tolerating minor defects or evolving quality standards—contrary to the common assumption that training data are perfectly normal. To bridge this gap, the study presents the first systematic formulation of this realistic setting, introduces tailored evaluation metrics, and proposes RePaste, a novel method that iteratively re-pastes image regions with high anomaly scores back into the input to adaptively enhance the model’s discrimination capability under ambiguous normality. Evaluated on a new benchmark derived from MVTec AD, RePaste achieves state-of-the-art performance under the proposed metrics while maintaining leading results in conventional AUROC and PRO scores, demonstrating its effectiveness and robustness in scenarios with ill-defined normal samples.
Current machine learning evaluation practices predominantly rely on surface-level performance metrics, often neglecting the internal mechanisms of models. This work proposes trustworthy interpretability as a central evaluation paradigm and, for the first time, systematically demonstrates that it satisfies core criteria from the philosophy of science—namely falsifiability, reproducibility, and predictive power. By constructing an evaluation framework that integrates causal analysis with mechanistic probing, the study delineates three functional pathways through which interpretability enables the identification of behavioral origins, detection of latent flaws, and prediction of potential failure modes. This approach advances model assessment beyond performance-oriented benchmarks toward a deeper understanding of underlying mechanisms.
This work addresses the limitation of existing safety-critical systems, which typically evaluate only predictive accuracy while lacking rigorous validation of the overall calibration of predicted probability distributions. To bridge this gap, the authors propose a modular calibration testing framework that decouples the calibration process into four interchangeable components: data model, scoring rule, hypothesis formulation, and statistical test procedure. Built upon formal statistical hypothesis testing, the framework provides a single accept/reject decision for the entire predictive distribution. Crucially, it rejects only overly confident predictions while tolerating reasonable deviations, thereby balancing practicality with flexibility. Empirical evaluations on weather forecasting and robotic pose estimation tasks demonstrate that the framework effectively supports reliable deployment in safety-critical applications.