distribution matching

Techniques to measure, align, and steer data distributions across datasets or modalities, including detecting distribution shift, generating samples that match target distributions, and applying mitigation or alignment procedures.

distributionmatching

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Existing dataset characterization methods—statistical, structural, and model-driven—lack sufficient interpretability and deep structural insight. To address this, we propose a novel tensor-based representation paradigm that transcends conventional two-dimensional assumptions. Our approach leverages high-order tensor decomposition, multilinear modeling, and cross-modal joint representation to explicitly capture high-dimensional, nonlinear, and multi-source relational structures inherent in complex data. Extensive experiments demonstrate that the proposed method significantly outperforms baseline approaches in three key aspects: (i) disentangling intricate data structures, (ii) enhancing feature interpretability, and (iii) enabling traceable downstream task reasoning. This work establishes a unified tensor modeling framework for dataset representation and pioneers a data-driven discovery pathway tailored for explainable AI. It contributes both theoretical advances—through formalizing multilinear structure learning—and practical utility—by providing an interpretable, computationally grounded toolkit for transparent data analysis.

Enhancing interpretability and intelligence in complex dataset analysisProposing tensor-based techniques for improved data understandingSurveying limitations of current dataset characterization methods

Methods for quantifying dataset similarity: a review, taxonomy and comparison

Dec 07, 2023
MS
Marieke Stolte
🏛️ TU Dortmund University

This study addresses the critical need for quantifying dataset similarity in model generalization, transfer learning, simulation calibration, and two-sample testing. We systematically survey 118 similarity quantification methods and propose the first ten-dimensional classification framework, organizing approaches into seven technical categories: statistical distances (e.g., Wasserstein distance, Maximum Mean Discrepancy), kernel-based methods, information-theoretic measures, dimensionality-reduction embeddings, permutation tests, generative-model-based discriminators, and Gaussian process likelihood ratios. We develop a multi-dimensional evaluation system balancing theoretical guarantees, interpretability, and practical applicability, yielding a structured recommendation matrix aligned with task requirements and data characteristics. Furthermore, we introduce the first open-source, interactive tool for method selection—enabling real-time filtering and parameter configuration—to significantly enhance both selection efficiency and deployment suitability.

Compare 118 methods on applicability and interpretabilityProvide recommendations for selecting dataset similarity measuresReview and classify methods for quantifying dataset similarity

Must-Read Papers

Most classic and influential ideas
View more

Existing methods struggle to align and interpret distribution shifts across heterogeneous, domain-consistent datasets—such as tabular, textual, visual, and time-series data—especially when scale and modality disparities are pronounced, resulting in poor interpretability. This paper introduces the first human-centric, cross-modal distribution discrepancy explanation framework, implemented as an interpretable dataset comparison toolbox. It integrates statistical hypothesis testing, feature importance decomposition, class activation mapping (CAM), contrastive representation learning, and interpretable generative modeling to enable fine-grained, semantically readable attribution and visualization of distributional shifts. Evaluated across diverse real-world scenarios, the framework significantly improves users’ efficiency in understanding shift causes and enhances the accuracy of intervention decisions—thereby overcoming the limitations of conventional black-box shift detection approaches.

Data InterpretationInter-data DifferentiationMulti-type Data Analysis

Multiple Distribution Shift -- Aerial (MDS-A): A Dataset for Test-Time Error Detection and Model Adaptation

Feb 18, 2025
NN
Noel Ngu
🏛️ Arizona State University | Universidad Nacional del Sur | U.S. Department of Defense | United States Military Academy

This work addresses the severe performance degradation and erroneous detection of aerial vision models under weather-induced distribution shifts. To this end, we introduce MDS-A—the first multi-distribution-shift benchmark for aerial imagery—featuring high-fidelity synthetic training data generated in Unreal Engine under six controlled meteorological conditions, a mixed-weather test set, and comprehensive annotations. We propose EDR (Error Detection and Recovery), a knowledge-driven framework enabling test-time uncertainty modeling and lightweight self-adaptation. MDS-A is the first benchmark to support fine-grained out-of-distribution (OOD) attribution analysis and standardized evaluation across multidimensional weather shifts. Experiments show that mainstream YOLOv5/v8 models suffer 32–68% mAP drops across weather domains; with EDR, erroneous detection accuracy reaches 89.7%, and online model adaptation is effectively triggered.

Addresses performance degradation due to distribution shiftsEvaluates models under varied simulated weather conditionsIntroduces MDS-A dataset for error detection and adaptation

Rethinking Distribution Shifts: Empirical Analysis and Inductive Modeling for Tabular Data

Jul 11, 2023
JL
Jiashuo Liu
🏛️ Tsinghua University | Columbia University

Distribution shift in tabular data lacks empirical grounding, and existing robust learning methods—such as distributionally robust optimization (DRO)—frequently fail in real-world settings. Prior work predominantly assumes covariate shift (X-shift), whereas empirical analysis reveals that label-conditional shift (Y|X-shift) is both more prevalent and dominant in practice. Method: The authors construct a comprehensive benchmark platform spanning five real-world tabular datasets and 60,000 experimental configurations, enabling large-scale ablation studies on robustness. Contribution/Results: (1) Theory-driven methods like DRO offer no consistent advantage over empirical risk minimization (ERM); (2) implementation details—including model selection, hyperparameter tuning, and imbalance handling—exert far greater influence on robustness than the design of ambiguity sets; (3) a data-driven, inductive paradigm should supplant prior structural assumptions. These findings challenge the conventional paradigm in distributionally robust learning and provide reproducible, empirically grounded guidance for co-optimizing algorithms and data.

Analyzing real-world distribution shifts in tabular datasetsEvaluating robust algorithms' performance against empirical shiftsIdentifying implementation factors affecting distributionally robust optimization

Automatic dataset shift identification to support root cause analysis of AI performance drift

Nov 12, 2024
MR
Mélanie Roschewitz
🏛️ Imperial College London

In AI deployment for medical imaging, data distribution shifts frequently cause abrupt performance degradation and increased misdiagnosis risk. Existing methods can only detect the presence of shift but fail to identify its specific type—e.g., covariate shift, prior (concept) shift, or compound shift—hindering root-cause analysis and targeted mitigation. This paper proposes the first unsupervised framework for data shift type identification. It introduces a novel joint shift detection mechanism that synergistically leverages self-supervised encoder representations and task-model outputs. By integrating feature distribution comparison, unsupervised clustering, and multimodal image modeling, the method achieves high-accuracy shift-type discrimination across three major imaging modalities—chest X-ray, mammography, and fundus photography—and five realistic shift scenarios. Evaluated on four large public medical imaging datasets, it significantly enhances the robustness and interpretability of clinical AI systems.

Distinguish prevalence, covariate, and mixed shifts unsupervisedIdentify diverse dataset shifts in medical imaging AIImprove shift detection using self-supervised encoders

Data fusion using weakly aligned sources

Aug 28, 2023
SL
Sijia Li
🏛️ University of Washington | Fred Hutchinson Cancer Center | Harvard T.H. Chan School of Public Health

Addressing the challenge of smooth finite-dimensional parameter estimation under weak alignment—where multi-source data exhibit partial,而非 perfect, correspondence and fully aligned samples are scarce—this paper proposes a novel semiparametric data fusion method. We establish, for the first time, the semiparametric efficiency bound under weak alignment and develop a theoretically grounded, robust estimator that jointly models alignment uncertainty and leverages auxiliary information, thereby substantially reducing reliance on fully aligned samples. Our approach relaxes the stringent strong-alignment assumption inherent in conventional fusion frameworks. Applied to an HIV monoclonal antibody prevention trial, it successfully quantifies the association between neutralizing antibodies and viral genotypes, demonstrating improved statistical efficiency and practical applicability. Key contributions include: (i) derivation of the semiparametric efficiency bound under weak alignment; (ii) a computationally feasible, robust fusion algorithm with provable efficiency; and (iii) interpretable, real-world validation in a clinical setting.

Addresses scarcity of fully aligned sources in data fusionEstimates smooth parameters using weakly aligned data sourcesQuantifies efficiency gains from integrating misaligned sources

Latest Papers

What's happening recently
View more

Shift is Good: Mismatched Data Mixing Improves Test Performance

Oct 28, 2025
MM
Marko Medvedev
🏛️ University of Chicago | Tsinghua University | Toyota Technological Institute at Chicago

This paper investigates distribution shift arising from mismatched training-to-test proportions across subpopulations—i.e., when the mixture proportions differ between training and test distributions—even in the absence of statistical dependencies or transferable structures among subpopulations. Method: We formalize the problem via mixture distribution modeling, conduct rigorous theoretical analysis, and derive closed-form solutions for the optimal training mixture proportions that maximize test performance. Contribution/Results: We identify and characterize the counterintuitive phenomenon of “beneficial distribution shift”: deliberately deviating from proportional sampling can significantly improve generalization. We establish tight bounds on the achievable performance gain and derive the optimal training proportions under diverse multi-scenario settings. Furthermore, we extend our framework to practical applications such as skill composition tasks. This work broadens the distribution shift research paradigm by providing interpretable, theoretically grounded principles for data proportion design—offering both conceptual insight and actionable guidelines for real-world deployment.

Extends analysis to compositional settings with varying skill distributionsIdentifies optimal training proportions for mismatched data mixturesInvestigates beneficial effects of distribution shift on test performance

MechDetect: Detecting Data-Dependent Errors

Dec 03, 2025
PJ
Philipp Jung
🏛️ Berlin University of Applied Sciences and Technology

The core challenge in data quality monitoring lies in error provenance—specifically, identifying the underlying mechanisms that generate errors—a problem largely overlooked by existing work, which seldom models such mechanisms explicitly. This paper focuses on errors arising from intrinsic dependencies within data and proposes MechDetect, the first method to systematically extend missing-data mechanism detection to diverse error types—including outliers, inconsistencies, and format violations. Leveraging joint statistical modeling and supervised learning, MechDetect simultaneously models tabular data and their error masks to automatically determine whether observed errors stem from inherent characteristics of the original data. Extensive experiments across multiple benchmark datasets demonstrate that MechDetect significantly outperforms state-of-the-art baselines in accurately diagnosing error-generation mechanisms. By providing mechanistic interpretability, it establishes a theoretical foundation and practical framework for explainable data repair.

Detect data-dependent error generation mechanismsEstimate error dependency using machine learning modelsExtend missing value analysis to other error types

Towards Mitigating Systematics in Large-Scale Surveys via Few-Shot Optimal Transport-Based Feature Alignment

Nov 14, 2025
SH
Sultan Hassan
🏛️ Space Telescope Science Institute | South African Radio Astronomy Observatory | University of the Western Cape | Johns Hopkins University

In large-scale astronomical surveys, systematic errors induce distributional shifts in observations, severely degrading the generalization of pre-trained models under label-scarce conditions. To address this, we propose an optimal transport (OT)-based few-shot feature alignment method that directly aligns in-distribution (ID) and out-of-distribution (OOD) sample distributions in the pre-trained feature space—without requiring paired samples or prior knowledge of systematic errors. Our approach jointly optimizes mean squared error and OT distance to achieve robust transfer learning. We validate the method on both MNIST and real-world neutral hydrogen (HI) large-scale sky maps. Results demonstrate substantial improvements in downstream task performance, particularly in realistic astronomical settings where systematic errors are difficult to model and labeled data are extremely scarce. The method exhibits strong adaptability and generalization capability under such challenging conditions.

Addressing distribution shifts when applying pre-trained models to new dataAligning feature distributions between simulated and real observational dataMitigating systematic errors in large-scale astronomical survey data

This study addresses the lack of systematic and impartial evaluation of existing methods for measuring distributional similarity in numerical data, which hinders informed selection in practice. The authors construct the first comprehensive benchmarking framework encompassing 36 similarity measures for continuous data—including statistical tests, distance-based metrics, and embedding approaches—and evaluate their discriminative power and computational efficiency through large-scale simulations across diverse distributional discrepancies (e.g., shifts in location, scale, and higher-order moments) and both two-sample and multi-sample settings. Based on empirical performance, the work proposes a data-characteristic-driven strategy for method selection, establishes a performance ranking, and demonstrates that combining only four to six methods suffices to achieve near-optimal performance in 90%–95% of scenarios.

dataset similaritydistribution comparisonk-sample testing

This study addresses systematic biases in large language models (LLMs) when simulating survey responses, including skewed marginal distributions, poor variance calibration, and attenuated variable relationships. It proposes a novel decomposition of simulation fidelity into three quantifiable dimensions—structural, marginal, and individual—and systematically evaluates the multidimensional fidelity of three mitigation strategies: prompt engineering, output post-processing, and few-shot fine-tuning, using small-scale pilot data. Empirical results demonstrate that few-shot fine-tuning achieves a favorable balance across fidelity dimensions; however, uneven fidelity across subpopulations may compromise the consistent alignment of diverse viewpoints.

LLM-based survey simulationpluralistic alignmentsmall pilot data

Hot Scholars

KC

Kim Christensen

Imperial College London
Complexity & Networks ScienceStatitical Physics
EN

Eric Nalisnick

Assistant Professor, Johns Hopkins University
Machine Learning
YW

Yang Weng

Associate Professor, School of Electrical, Computer, and Energy Eng., Arizona State University
Machine Learning for Power Systems
AL

Anqi Liu

Tulane University
Human GeneticsComputational BiologyBioinformaticsDeep Learning
MP

Mayukha Pal

Global R&D Leader - Cloud & Advanced Analytics, ABB Ability Innovation Center
Data SciencePhysics-Aware AnalyticsPower System AnalyticsBiomedical Signal Processing