distribution comparison

Selecting and applying statistical divergences or metrics to compare probability distributions, produce interpretable closed-form measures of agreement or change, and quantify how actions alter a downstream evaluator's distribution over candidate outcomes.

distributioncomparison

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the lack of systematic investigation into the statistical properties and testing performance of Jensen–Shannon divergence (JSD) and Kullback–Leibler (KL) divergence in credit risk model monitoring. It derives, for the first time, chi-squared asymptotic reference distributions for both divergences under distributional shift using asymptotic theory, and conducts a comprehensive Monte Carlo simulation to evaluate their Type I error control and statistical power relative to the Population Stability Index (PSI). The results demonstrate that JSD exhibits superior Type I error control, closely attaining the nominal 5% level, yet shows limited power (27%) in small samples (n=200); in contrast, KL divergence and PSI achieve higher power (32%). These findings provide both theoretical grounding and empirical guidance for selecting appropriate divergence metrics in practical model monitoring.

credit riskdistributional shiftdivergence measures

This work addresses the lack of reliable statistical evaluation methods for generative models, which hinders the assessment of their generalization performance and the estimability of evaluation metrics from finite samples. The authors propose a theoretical framework that systematically analyzes the conditions under which common evaluation metrics are statistically estimable, distinguishing between test-class-based metrics and divergence-based metrics in finite-sample settings. Leveraging tools from integral probability metrics (IPMs), Rényi divergences, and fat-shattering dimension, they rigorously establish—for the first time—that IPMs induced by bounded test classes admit arbitrarily accurate estimation from finite samples, whereas KL and Rényi divergences, which depend on rare events, do not. This study provides a foundational theoretical basis and practical guidance for evaluating generative models.

evaluabilityfinite samplesgenerative models

Diffusion-Based Hypothesis Testing and Change-Point Detection

Jun 19, 2025
SM
Sean Moushegian
🏛️ Duke University | University of Pittsburgh

To address the limited statistical power and poor changepoint localization accuracy of score-based tests in likelihood-free inference, this paper proposes a novel hypothesis testing and changepoint detection framework grounded in diffusion divergence. The key innovation lies in the first integration of score functions into the diffusion divergence paradigm, augmented by a learnable weighted matrix that modulates the score function to enhance discriminative capability. Theoretically, we derive a tight performance bound on detection power and establish sufficient conditions for achieving optimality. Methodologically, we design a numerical optimization algorithm for the weighting matrix and a statistically principled stopping rule. Monte Carlo experiments demonstrate that, compared to conventional score-based methods, the proposed approach reduces changepoint localization error by 32% and improves test power by over 18%, substantially narrowing the performance gap with likelihood-based approaches.

Extend score-based hypothesis testing to diffusion-based methodsOptimize weight matrix for improved detection performanceTheoretically analyze diffusion-based algorithms' optimal scenarios

This work addresses the critical challenge of accurately modeling preference functions that aggregate multidimensional criteria into holistic judgments in settings such as admissions and medical diagnosis. Departing from conventional assumptions of linearity or strong structural forms, the paper proposes the first robust nonparametric learning algorithm that achieves optimal performance without requiring any prior knowledge of the preference structure, assuming only monotonic non-decreasing behavior across each criterion. Theoretical analysis demonstrates the severe consequences of common model misspecifications, while experiments on both synthetic and real-world data confirm that the method maintains statistical efficiency under linear preferences and reliably recovers true evaluator preferences in general cases. Notably, the approach effectively uncovers key behavioral differences between human evaluators and large language models in their assessment strategies.

evaluator preferencesmodel mismatchmulti-criteria evaluation

Existing methods struggle to align and interpret distribution shifts across heterogeneous, domain-consistent datasets—such as tabular, textual, visual, and time-series data—especially when scale and modality disparities are pronounced, resulting in poor interpretability. This paper introduces the first human-centric, cross-modal distribution discrepancy explanation framework, implemented as an interpretable dataset comparison toolbox. It integrates statistical hypothesis testing, feature importance decomposition, class activation mapping (CAM), contrastive representation learning, and interpretable generative modeling to enable fine-grained, semantically readable attribution and visualization of distributional shifts. Evaluated across diverse real-world scenarios, the framework significantly improves users’ efficiency in understanding shift causes and enhances the accuracy of intervention decisions—thereby overcoming the limitations of conventional black-box shift detection approaches.

Data InterpretationInter-data DifferentiationMulti-type Data Analysis

Latest Papers

What's happening recently
View more

This study addresses the reliability and validity of LLM-as-judge evaluations, which are susceptible to shifts in the judge model’s version even when candidate responses remain unchanged. The authors conduct a systematic audit of dense Qwen3 models (1.7B–32B) and MiniMax API iterations (M2 to M2.7) across four benchmark judgment datasets. They propose a multidimensional auditing framework incorporating multiscale judge comparisons, repeated-sampling juries, structured debate protocols, and probes for position and verbosity biases. Findings indicate that only the upgrade from Qwen3-1.7B to -4B yields consistent performance gains; stronger judges mitigate but do not eliminate systematic biases; and structured debate substantially alters verdicts, though reliable attribution requires access to detailed interaction logs.

evaluator biasjudge reliabilityLLM-as-judge

This work addresses the instability of preference judgments from proprietary large language model (LLM) evaluators, which are highly sensitive to version updates and thus yield unreliable short-term assessments. To systematically characterize evaluator-driven preference dynamics, the authors propose EPC, a diagnostic framework that integrates the Multimodal Preference Collapse Index (MPCI), evaluator coupling matrices, Jensen–Shannon divergence, and output-format confusion analysis. Experiments reveal significant fluctuations in preference coupling and self-evaluation mechanism collapse across versions of models such as GPT-4o, demonstrating the severe limitations of single-snapshot evaluations. The study releases all data and tools publicly, establishing a new paradigm for robust and reliable LLM evaluation.

diagnostic frameworkevaluator instabilitypreference collapse

Current machine learning evaluation practices predominantly rely on surface-level performance metrics, often neglecting the internal mechanisms of models. This work proposes trustworthy interpretability as a central evaluation paradigm and, for the first time, systematically demonstrates that it satisfies core criteria from the philosophy of science—namely falsifiability, reproducibility, and predictive power. By constructing an evaluation framework that integrates causal analysis with mechanistic probing, the study delineates three functional pathways through which interpretability enables the identification of behavioral origins, detection of latent flaws, and prediction of potential failure modes. This approach advances model assessment beyond performance-oriented benchmarks toward a deeper understanding of underlying mechanisms.

behavioral metricsinterpretabilitymachine learning

This study addresses sample bias arising from unequal probability sampling or unobservable events by proposing a novel statistical test based on Rényi divergence, which is introduced here for the first time in the context of bias detection. The authors construct test statistics applicable to both uncensored and Type-I censored data, establishing favorable theoretical properties such as asymptotic normality. Critical values are determined via Monte Carlo simulations, and the method demonstrates robust performance across varying sample sizes and censoring proportions. Empirical analysis of two real-world datasets successfully identifies length bias, with the proposed test exhibiting higher power than existing approaches based on Kullback–Leibler divergence and likelihood ratios.

biased samplelength biasRenyi divergence

This study identifies and formally names the “Metric Aggregation Discrepancy” (MAD) problem, wherein inconsistent metric aggregation across pipeline stages in coupled agent-based modeling and multi-objective evolutionary algorithm systems leads to erroneous policy recommendations and irreproducible results. To address this, the authors propose a “metric contract” mechanism that enforces a unified metric extraction interface across the entire pipeline at the architectural level, ensuring computational consistency during scheduling. The approach incurs only approximately 3% runtime overhead while substantially improving reliability: in reproducing EpidemiOptim, it reduces policy recommendation error rates from 83% to near zero, and in the Lake Problem benchmark, it increases joint threshold success rates from 0.401 to 0.552.

Agent-Based ModelingMetric Aggregation DivergenceMulti-Objective Evolutionary Algorithm

Hot Scholars

HP

Henri Prade

CNRS, France and University of New South Wales
Artificial IntelligenceDecision Making
XL

Xing Liu

Imperial College London; QuantCo
Computational StatisticsKernel MethodsHypothesis testing
YK

Yongdai Kim

Seoul National University
statisticsmachine learning