optimal transport divergence

Designs and implements methods to compute and analyze optimal-transport divergences between probability distributions or empirical samples — e.g., Earth Mover’s Distance / Wasserstein distance — by specifying ground-costs, building cost matrices, and solving for transport plans using exact linear-program solvers or entropic-regularized approximations (e.g., Sinkhorn) with attention to numerical efficiency and differentiability. Uses these computed distances as evaluation or training quantities to measure fidelity and diversity, produce divergence-based scores, and detect mode-coverage or distributional mismatches.

optimaltransportdivergence

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.2
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Multi-Dimensional Wasserstein Distance Implementation in Scipy

Oct 25, 2025
ZL
Zehao Lu
🏛️ Utrecht University

SciPy currently supports only one-dimensional Wasserstein distance computation, lacking native capability to model optimal transport distances between multidimensional probability distributions. Method: This work introduces the first implementation of a multidimensional Wasserstein distance module in SciPy, formulated as a linear programming (LP) problem in standard form and solved efficiently using SciPy’s built-in `linprog` solver. The implementation handles discrete distributions of arbitrary dimensionality, supports batched inputs and sparse optimization, and includes comprehensive unit tests, documentation, and usage examples. Contribution/Results: The code has been merged into SciPy’s main branch and will be released in version 1.13+, filling a critical gap in SciPy’s support for optimal transport and multidimensional statistical distances. This advancement significantly broadens SciPy’s applicability in machine learning, bioinformatics, and computational statistics, enabling efficient, library-native computation of high-dimensional optimal transport metrics.

Enhances Scipy capabilities for multi-dimensional statistical analysisExtends Wasserstein distance to multi-dimensional distributions in ScipyTransforms optimal transport problem into linear programming formulation

Synthesis and Analysis of Data as Probability Measures with Entropy-Regularized Optimal Transport

Jan 13, 2025
BM
Brendan Mallery
🏛️ Tufts University | The NSF AI Institute for Artificial Intelligence and Fundamental Interactions

This work addresses the problem of probabilistic distribution modeling and robust classification for point cloud data under few-shot learning settings. We propose a novel framework unifying probabilistic measure synthesis (via weighted Wasserstein-2 barycenters) and analysis (via barycentric coordinate estimation), regularized by entropy. For the first time under minimal assumptions, we derive its gradient analytically and construct a convex quadratic programming solver based on the entropic map fixed-point equation. We establish dimension-free convergence rates for barycentric coordinates and Wasserstein stability guarantees. Integrating Sinkhorn divergence, entropic maps, and optimal transport, our method significantly improves recognition accuracy for corrupted point cloud classes in few-shot scenarios—outperforming neural network baselines. Extensive experiments validate both theoretical convergence and robustness to noise.

Data BiasOutlier Detection in Point CloudsWeighted Probability Estimation

Shift-Dispersion Decompositions of Wasserstein and Cram'er Distances

Aug 19, 2024
JR
Johannes Resin
🏛️ Goethe University Frankfurt | Heidelberg Institute for Theoretical Studies | Karlsruhe Institute of Technology

Wasserstein and Cramér distances lack directional interpretability—i.e., they do not distinguish between location shifts and scale changes in distributional discrepancies. To address this, we propose an interpretable framework based on geometric decomposition of quantile functions, uniquely disentangling each distance into directed shift (location) and dispersion (scale) components. This decomposition is the first to satisfy naturalness, uniqueness, and additivity—key statistical desiderata—within the location-scale family. We further derive explicit sensitivity expressions of the distances with respect to location and dispersion parameters, and establish a weak stochastic order theory that jointly characterizes both location and dispersion orderings. Empirically validated on extreme temperature forecasting evaluation and probabilistic survey design in economics, our method substantially enhances semantic clarity in interpreting distributional differences and strengthens decision-support capabilities.

Apply decompositions to forecast evaluation and probabilistic survey designDecompose Wasserstein and Cramér distances into shift and dispersion componentsEnhance interpretability of divergence measures between probability distributions

Sinkhorn Distributionally Robust Optimization

Sep 24, 2021
JW
Jie Wang
🏛️ Georgia Institute of Technology | University of Texas at Austin

This work addresses key limitations of Wasserstein distributionally robust optimization (DRO), including high computational cost in computing the Wasserstein distance and the unrealistic nature of the worst-case distribution. We propose a general DRO framework based on the Sinkhorn distance—an entropy-regularized variant of the Wasserstein distance. First, we establish a convex dual characterization of Sinkhorn-DRO applicable to arbitrary nominal distributions, transportation costs, and loss functions, ensuring both tractability and modeling flexibility while yielding more realistic worst-case distributions. Second, we design a stochastic mirror descent algorithm with bias-corrected gradient estimation, achieving near-optimal sample complexity and supporting both smooth and nonsmooth losses. Experiments on synthetic and real-world datasets demonstrate that our method significantly improves computational efficiency, generalization performance, and robustness compared to classical Wasserstein DRO.

Derive convex dual reformulation for general distributionsDevelop stochastic mirror descent algorithm for solutionStudy distributionally robust optimization with Sinkhorn distance

This paper investigates the stability and exponential convergence of the Sinkhorn algorithm for entropy-regularized optimal transport. Focusing on quadratic cost, it establishes a Wasserstein stability theory based on semi-concavity assumptions, yielding the first global exponential convergence guarantee under log-concave marginals. It derives sharp convergence rate bounds with linear dependence on the regularization parameter. The analysis is extended to novel non-compact, unbounded settings—including Riemannian manifolds, elastic costs, and light-tailed marginals. Methodologically, the work integrates semi-concavity analysis, uniform upper bounds on the Hessian of Sinkhorn potentials, and refined Wasserstein distance estimation. Collectively, this provides the first unified exponential convergence guarantee for a broad class of generalized cost functions and marginal distributions. The derived rates improve upon prior results, significantly expanding the theoretical applicability of the Sinkhorn algorithm.

Establishing exponential convergence under semiconcavity without bounded costExtending convergence results to various marginals and cost structuresStudying stability of entropic optimal transport optimizers and Sinkhorn convergence

Latest Papers

What's happening recently
View more

Optimal transport (OT) is a central framework for modeling distribution shifts. Because OT compares distributions directly in input space, a well-designed ground metric between observations is essential to ensure that the optimizer does not violate the true geometry of change. We propose Displacement-Reshaped Optimal Transport (ReshapeOT), a method that reshapes the ground metric by integrating observed sample displacements as an additional source of knowledge. Technically, ReshapeOT replaces the Euclidean metric with a Mahalanobis distance estimated from displacement second moments. This effectively carves expressways through the input space, inviting transport solutions that better align with observed displacements. Our method is computationally lightweight, integrates seamlessly into any OT solver that operates on a cost matrix, and can be kernelized for further flexibility. Experiments on synthetic and real-world data show that ReshapeOT achieves substantial gains in transport reliability. We further demonstrate our method's usefulness in two practical use cases.

distribution shiftsground metricoptimal transport

This work addresses the limitations of existing convergence bounds for the Sinkhorn–Knopp algorithm in the presence of outliers, which heavily depend on the regularization parameter or element-wise ratios and thus poorly reflect practical performance. To overcome this, we introduce the notion of “well-boundedness” to characterize the intrinsic quality of the dominant data structure and combine it with a pre-scaling technique to effectively isolate the influence of outliers. Building on this framework, we uncover a density-threshold-driven phase transition phenomenon in matrix scaling and establish a novel convergence analysis. Under the well-boundedness condition, the algorithm achieves ε-accuracy in only O(log(1/ε)) iterations, providing the first rigorous convergence guarantee that is independent of problem dimension, regularization cost, and outlier contamination.

entropically regularized optimal transportiteration complexitymatrix scaling

This work addresses the limitation of traditional causal inference methods—such as average treatment effects—which capture only local differences in outcome distributions and thus fail to fully characterize the treatment’s impact on the entire distribution. The authors propose the Sinkhorn Treatment Effect, which for the first time integrates entropy-regularized optimal transport into causal inference. By constructing a smooth transformation of counterfactual mean embeddings, they derive a differentiable functional representation of distributional treatment effects. Building on this framework, they develop a debiased estimator with asymptotic efficiency and a multi-regularization-parameter aggregation test. Both theoretical analysis and empirical experiments demonstrate that the proposed approach substantially enhances the identification and detection of distributional causal effects on synthetic and image data.

causal inferencecounterfactual distributionsdistributional divergence

This work proposes a class of structure-aware divergences that explicitly incorporate geometric relationships among elements in the support set of probability distributions—addressing a key limitation of classical information-theoretic measures such as Shannon entropy and f-divergences, which disregard structural similarities. By integrating the underlying geometry of the support set, the authors define a structure-aware entropy and derive corresponding Bregman divergences that retain desirable properties of the Kullback–Leibler divergence and Shannon entropy while embedding pairwise similarities directly into the divergence formulation. The approach successfully uncovers structural patterns missed by conventional methods in synthetic clustering tasks, achieves computational efficiency several orders of magnitude higher than optimal transport, and reproduces and extends established findings in applications to economic geography and ecology.

f-divergencesprobability distributionsShannon entropy

Hot Scholars

NC

Ningyuan Chen

Department of Management, UTM & Rotman School of Management, University of Toronto
Revenue ManagementOnline LearningOperations ManagementBusiness Analytics
RB

Ricardo Baeza-Yates

KTH - Univ. Pompeu Fabra - Univ. de Chile; Sweden, Spain & Chile
Responsible AIInformation retrievalWeb searchWeb mining
LZ

Liangyu Zhang

assistant professor at Shanghai University of Finance and Economics
reinforcement learningstatistical learning theory
YD

Yuexi Du

PhD candidate @ Yale University
Computer VisionMedical Image AnalysisMulti-modal Learning
RP

Rita P. Ribeiro

Faculty of Sciences, University of Porto and INESC TEC
Imbalanced Domain LearningAnomaly DetectionExplainable AIAI for Social Good