analyze prediction discrepancies

Designs and implements methods to measure, quantify, and interpret differences between predictions produced by multiple models, modalities, or model versions, including construction of discrepancy metrics and tests that summarize prediction gaps. Builds procedures that use these discrepancy measures to detect anomalous inputs or distribution shifts (e.g., derive OOD or risk scores, flag inputs with large prediction gaps) and evaluates and validates those signals on held-out data.

analyzepredictiondiscrepancies

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Tests for model misspecification in simulation-based inference: from local distortions to global model checks

Dec 19, 2024
NA
Noemi Anau Montel
🏛️ Max-Planck-Institut für Astrophysik | Kavli Institute for Cosmology Cambridge | University of Cambridge | GRAPPA Institute | Institute for Theoretical Physics Amsterdam | University of Amsterdam

This work addresses the lack of systematic model misspecification diagnostics in simulation-based inference (SBI). We propose the first unified framework for multi-scale misspecification diagnosis. Methodologically, we develop a distortion-based statistical testing theory explicitly grounded in classical hypothesis testing; design a self-calibrating neural density estimation algorithm for end-to-end misspecification identification; and integrate distortion-driven testing, simulation-based Bayesian inference, and joint residual–outlier analysis. Our key contributions are: (i) the first interpretable and scalable testing paradigm bridging local anomaly detection to global model validation; and (ii) empirical validation across diverse simulation tasks and on the real gravitational-wave event GW150914—reproducing established results while successfully extending diagnostics to high-dimensional, complex forward models.

Applies the framework to real data, including gravitational wave analysis.Develops a simulation-based framework for model misspecification analysis.Introduces distortion-driven tests for detecting model discrepancies.

Detecting Conflicts in Evidence Synthesis Models Using Score Discrepancies

Nov 04, 2025
FY
Fuming Yang
🏛️ University of Cambridge | National University of Singapore

This study addresses structural conflicts—arising both among heterogeneous data sources and between data and model assumptions—in evidence synthesis models. We propose a general conflict detection framework based on score-based discrepancy measures. Methodologically, we extend prior–data conflict diagnostics to the latent space of hierarchical models, enabling inconsistency detection under multilevel and non-exchangeable structures; integrating Bayesian evidence synthesis, score-function-based metrics, and posterior simulation, our approach provides quantitative assessment of model assumption–data compatibility. Key contributions include: (1) moving beyond conventional bias diagnostics confined to the prior–likelihood level; (2) demonstrating high sensitivity to conflicts in both exchangeable and non-exchangeable models; and (3) exhibiting complementary diagnostic capability to existing methods in a real-world influenza severity model, thereby significantly enhancing the reliability of complex Bayesian inference.

Detecting conflicts in evidence synthesis modelsExtending conflict diagnostics to hierarchical modelsQuantifying inconsistencies between data sources

Estimating and evaluating counterfactual prediction models

Aug 24, 2023
CB
Christopher B. Boyer
🏛️ Cleveland Clinic Research | Case Western Reserve University | Harvard T.H. Chan School of Public Health | Richard A. and Susan F. Smith Center for Outcomes Research | Beth Israel Deaconess Medical Center | Brown University School of Public Health

Counterfactual prediction under evolving intervention policies or hypothetical decision scenarios remains challenging due to unobservable potential outcomes, hindering model identifiability, evaluation, and generalization. Method: We propose the first systematic theoretical framework addressing this challenge—comprising (i) identifiability conditions for counterfactual prediction models, (ii) a performance evaluation system targeting loss, AUC, and calibration, and (iii) robust hyperparameter selection under model misspecification. Our approach integrates causal inference principles, doubly robust estimation, and loss-driven evaluation metric design. Contribution/Results: Validated via simulation studies and a real-world clinical application—cardiovascular risk prediction in statin-naïve populations—the framework significantly improves out-of-distribution generalization and clinical decision reliability in counterfactual settings.

Estimating counterfactual prediction models under different treatment policiesEvaluating model performance without observed potential outcomesProviding valid performance estimates under model misspecification

Existing methods struggle to align and interpret distribution shifts across heterogeneous, domain-consistent datasets—such as tabular, textual, visual, and time-series data—especially when scale and modality disparities are pronounced, resulting in poor interpretability. This paper introduces the first human-centric, cross-modal distribution discrepancy explanation framework, implemented as an interpretable dataset comparison toolbox. It integrates statistical hypothesis testing, feature importance decomposition, class activation mapping (CAM), contrastive representation learning, and interpretable generative modeling to enable fine-grained, semantically readable attribution and visualization of distributional shifts. Evaluated across diverse real-world scenarios, the framework significantly improves users’ efficiency in understanding shift causes and enhances the accuracy of intervention decisions—thereby overcoming the limitations of conventional black-box shift detection approaches.

Data InterpretationInter-data DifferentiationMulti-type Data Analysis

Handling Missingness, Failures, and Non-Convergence in Simulation Studies: A Review of Current Practices and Recommendations

Sep 27, 2024
SP
Samuel Pawel
🏛️ University of Zurich | University of Amsterdam | University of Marburg | Leiden University Medical Center | Ernst-Abbe University of Applied Sciences Jena

Accurately evaluating the performance of analytical methods in simulation studies is hindered by underreporting and inconsistent handling of “missingness” issues—such as algorithm failure or non-convergence—that compromise validity and reproducibility. Method: We conducted a large-scale empirical analysis of 482 methodological simulation studies, systematically extracting metadata, applying qualitative coding, and performing case studies—including publication bias correction—to quantify the prevalence and reporting practices of missingness. Contribution/Results: We found that only 23% of studies mentioned missingness and merely 14% described mitigation strategies. Based on these findings, we developed a novel missingness taxonomy tailored to simulation research and proposed actionable principles—including mandatory missingness reporting—alongside a comprehensive, end-to-end practice guideline covering reporting, handling, and replication. Validation confirmed substantial improvements in transparency, comparability, and reproducibility of simulation studies.

Addressing missing data in simulation studies due to failures.Providing solutions for non-convergence in methodological research.Recommending practices to report and handle simulation missingness.

Latest Papers

What's happening recently
View more

Current evaluation practices for supervised learning models are often misleading due to an overreliance on single aggregate metrics, which neglect the alignment among data characteristics, task objectives, and real-world application contexts. This work reframes model evaluation as a context-dependent, decision-oriented process and systematically investigates—through controlled experiments—the impact of dataset properties, validation strategies, class imbalance, and asymmetric error costs on evaluation outcomes. Leveraging diverse benchmark datasets, multiple validation protocols, and multidimensional performance measures, the study uncovers common pitfalls such as the accuracy paradox, data leakage, and metric misuse. It proposes a structured evaluation framework explicitly aligned with operational goals, offering principled guidance for developing more robust, reliable, and trustworthy supervised learning systems.

class imbalancemodel evaluationperformance metrics

This study addresses the problem of determining whether high-frequency monitoring data return to their pre-intervention baseline distribution following an intervention. The authors propose a sequential testing procedure that requires no assumptions about the underlying data distribution. The method constructs a discrepancy measure via universal inference and combines it with individualized empirical calibration to form a non-negative supermartingale, yielding an e-process that enables valid detection of the recovery time at any arbitrary stopping point without specifying a null model. Theoretical analysis provides finite-sample bounds on the calibration error, and both simulations and a clinical case study demonstrate the method’s superior performance in accurately identifying the time at which baseline conditions are restored.

distributional realignmenthigh-frequency monitoringintervention effect

This work addresses the practical challenge in industrial anomaly detection where “normal” samples often exhibit ambiguous definitions—such as tolerating minor defects or evolving quality standards—contrary to the common assumption that training data are perfectly normal. To bridge this gap, the study presents the first systematic formulation of this realistic setting, introduces tailored evaluation metrics, and proposes RePaste, a novel method that iteratively re-pastes image regions with high anomaly scores back into the input to adaptively enhance the model’s discrimination capability under ambiguous normality. Evaluated on a new benchmark derived from MVTec AD, RePaste achieves state-of-the-art performance under the proposed metrics while maintaining leading results in conventional AUROC and PRO scores, demonstrating its effectiveness and robustness in scenarios with ill-defined normal samples.

anomaly detectionevaluation metricsindustrial inspection

Current machine learning evaluation practices predominantly rely on surface-level performance metrics, often neglecting the internal mechanisms of models. This work proposes trustworthy interpretability as a central evaluation paradigm and, for the first time, systematically demonstrates that it satisfies core criteria from the philosophy of science—namely falsifiability, reproducibility, and predictive power. By constructing an evaluation framework that integrates causal analysis with mechanistic probing, the study delineates three functional pathways through which interpretability enables the identification of behavioral origins, detection of latent flaws, and prediction of potential failure modes. This approach advances model assessment beyond performance-oriented benchmarks toward a deeper understanding of underlying mechanisms.

behavioral metricsinterpretabilitymachine learning

This work addresses the limitation of existing safety-critical systems, which typically evaluate only predictive accuracy while lacking rigorous validation of the overall calibration of predicted probability distributions. To bridge this gap, the authors propose a modular calibration testing framework that decouples the calibration process into four interchangeable components: data model, scoring rule, hypothesis formulation, and statistical test procedure. Built upon formal statistical hypothesis testing, the framework provides a single accept/reject decision for the entire predictive distribution. Crucially, it rejects only overly confident predictions while tolerating reasonable deviations, thereby balancing practicality with flexibility. Empirical evaluations on weather forecasting and robotic pose estimation tasks demonstrate that the framework effectively supports reliable deployment in safety-critical applications.

calibrationdistributional validationprobabilistic forecasting

Hot Scholars

KO

Klaus Obermayer

Professor, Fakultät IV - Electrical Engineering and Computer Science, Technische Universität Berlin
Computational Neuroscience and Machine Learning
CC

Chen-Chen Zong

Nanjing University of Aeronautics & Astronautics
machine learning
RT

Ran Tian

Google Deepmind
Natural Language ProcessingMachine Learning
OH

Olaf Hellwich

Professor of Computer Vision & Remote Sensing, TU Berlin
Automatic Image Understanding3D Surface ReconstructionSynthetic Aperture Radar
TM

Tulika Mitra

Professor of Computer Science, National University of Singapore
Design AutomationLow Power DesignEmbedded SystemsReal-Time Systems