Score
Designs and implements methods to compute and analyze optimal-transport divergences between probability distributions or empirical samples — e.g., Earth Mover’s Distance / Wasserstein distance — by specifying ground-costs, building cost matrices, and solving for transport plans using exact linear-program solvers or entropic-regularized approximations (e.g., Sinkhorn) with attention to numerical efficiency and differentiability. Uses these computed distances as evaluation or training quantities to measure fidelity and diversity, produce divergence-based scores, and detect mode-coverage or distributional mismatches.
Optimal transport (OT) faces critical scalability, robustness, and ethical challenges in large-scale data settings. Method: This work systematically surveys OT’s evolution—from the Monge–Kantorovich foundation to modern computational approaches, including Sinkhorn iterations, primal-dual optimization, dimensionality reduction, and problem reduction techniques—and innovatively unifies scalability enhancements with emerging variants, notably Optimal Transport Warping (OTW). Contribution/Results: We present the first theoretically rigorous yet broadly applicable OT framework, achieving cross-domain practicality without sacrificing mathematical soundness. Empirical evaluation demonstrates that OTW significantly outperforms Dynamic Time Warping (DTW) in temporal alignment tasks. Our comprehensive OT algorithmic landscape establishes a new paradigm for high-dimensional distribution comparison—characterized by efficiency, numerical stability, and interpretability. Furthermore, the work proactively identifies key open challenges, including robust statistical modeling under distributional shift and fairness-aware constraints in OT-based inference.
SciPy currently supports only one-dimensional Wasserstein distance computation, lacking native capability to model optimal transport distances between multidimensional probability distributions. Method: This work introduces the first implementation of a multidimensional Wasserstein distance module in SciPy, formulated as a linear programming (LP) problem in standard form and solved efficiently using SciPy’s built-in `linprog` solver. The implementation handles discrete distributions of arbitrary dimensionality, supports batched inputs and sparse optimization, and includes comprehensive unit tests, documentation, and usage examples. Contribution/Results: The code has been merged into SciPy’s main branch and will be released in version 1.13+, filling a critical gap in SciPy’s support for optimal transport and multidimensional statistical distances. This advancement significantly broadens SciPy’s applicability in machine learning, bioinformatics, and computational statistics, enabling efficient, library-native computation of high-dimensional optimal transport metrics.
This work addresses the problem of probabilistic distribution modeling and robust classification for point cloud data under few-shot learning settings. We propose a novel framework unifying probabilistic measure synthesis (via weighted Wasserstein-2 barycenters) and analysis (via barycentric coordinate estimation), regularized by entropy. For the first time under minimal assumptions, we derive its gradient analytically and construct a convex quadratic programming solver based on the entropic map fixed-point equation. We establish dimension-free convergence rates for barycentric coordinates and Wasserstein stability guarantees. Integrating Sinkhorn divergence, entropic maps, and optimal transport, our method significantly improves recognition accuracy for corrupted point cloud classes in few-shot scenarios—outperforming neural network baselines. Extensive experiments validate both theoretical convergence and robustness to noise.
Wasserstein and Cramér distances lack directional interpretability—i.e., they do not distinguish between location shifts and scale changes in distributional discrepancies. To address this, we propose an interpretable framework based on geometric decomposition of quantile functions, uniquely disentangling each distance into directed shift (location) and dispersion (scale) components. This decomposition is the first to satisfy naturalness, uniqueness, and additivity—key statistical desiderata—within the location-scale family. We further derive explicit sensitivity expressions of the distances with respect to location and dispersion parameters, and establish a weak stochastic order theory that jointly characterizes both location and dispersion orderings. Empirically validated on extreme temperature forecasting evaluation and probabilistic survey design in economics, our method substantially enhances semantic clarity in interpreting distributional differences and strengthens decision-support capabilities.
This work addresses key limitations of Wasserstein distributionally robust optimization (DRO), including high computational cost in computing the Wasserstein distance and the unrealistic nature of the worst-case distribution. We propose a general DRO framework based on the Sinkhorn distance—an entropy-regularized variant of the Wasserstein distance. First, we establish a convex dual characterization of Sinkhorn-DRO applicable to arbitrary nominal distributions, transportation costs, and loss functions, ensuring both tractability and modeling flexibility while yielding more realistic worst-case distributions. Second, we design a stochastic mirror descent algorithm with bias-corrected gradient estimation, achieving near-optimal sample complexity and supporting both smooth and nonsmooth losses. Experiments on synthetic and real-world datasets demonstrate that our method significantly improves computational efficiency, generalization performance, and robustness compared to classical Wasserstein DRO.
This paper investigates the stability and exponential convergence of the Sinkhorn algorithm for entropy-regularized optimal transport. Focusing on quadratic cost, it establishes a Wasserstein stability theory based on semi-concavity assumptions, yielding the first global exponential convergence guarantee under log-concave marginals. It derives sharp convergence rate bounds with linear dependence on the regularization parameter. The analysis is extended to novel non-compact, unbounded settings—including Riemannian manifolds, elastic costs, and light-tailed marginals. Methodologically, the work integrates semi-concavity analysis, uniform upper bounds on the Hessian of Sinkhorn potentials, and refined Wasserstein distance estimation. Collectively, this provides the first unified exponential convergence guarantee for a broad class of generalized cost functions and marginal distributions. The derived rates improve upon prior results, significantly expanding the theoretical applicability of the Sinkhorn algorithm.
Optimal transport (OT) is a central framework for modeling distribution shifts. Because OT compares distributions directly in input space, a well-designed ground metric between observations is essential to ensure that the optimizer does not violate the true geometry of change. We propose Displacement-Reshaped Optimal Transport (ReshapeOT), a method that reshapes the ground metric by integrating observed sample displacements as an additional source of knowledge. Technically, ReshapeOT replaces the Euclidean metric with a Mahalanobis distance estimated from displacement second moments. This effectively carves expressways through the input space, inviting transport solutions that better align with observed displacements. Our method is computationally lightweight, integrates seamlessly into any OT solver that operates on a cost matrix, and can be kernelized for further flexibility. Experiments on synthetic and real-world data show that ReshapeOT achieves substantial gains in transport reliability. We further demonstrate our method's usefulness in two practical use cases.
This work addresses the limitations of existing convergence bounds for the Sinkhorn–Knopp algorithm in the presence of outliers, which heavily depend on the regularization parameter or element-wise ratios and thus poorly reflect practical performance. To overcome this, we introduce the notion of “well-boundedness” to characterize the intrinsic quality of the dominant data structure and combine it with a pre-scaling technique to effectively isolate the influence of outliers. Building on this framework, we uncover a density-threshold-driven phase transition phenomenon in matrix scaling and establish a novel convergence analysis. Under the well-boundedness condition, the algorithm achieves ε-accuracy in only O(log(1/ε)) iterations, providing the first rigorous convergence guarantee that is independent of problem dimension, regularization cost, and outlier contamination.
本文提出SinkSLOT方法,通过稀疏提升的运输计划解决大规模数据集上熵最优传输计算效率低和独立耦合问题。
This work addresses the limitation of traditional causal inference methods—such as average treatment effects—which capture only local differences in outcome distributions and thus fail to fully characterize the treatment’s impact on the entire distribution. The authors propose the Sinkhorn Treatment Effect, which for the first time integrates entropy-regularized optimal transport into causal inference. By constructing a smooth transformation of counterfactual mean embeddings, they derive a differentiable functional representation of distributional treatment effects. Building on this framework, they develop a debiased estimator with asymptotic efficiency and a multi-regularization-parameter aggregation test. Both theoretical analysis and empirical experiments demonstrate that the proposed approach substantially enhances the identification and detection of distributional causal effects on synthetic and image data.
This work proposes a class of structure-aware divergences that explicitly incorporate geometric relationships among elements in the support set of probability distributions—addressing a key limitation of classical information-theoretic measures such as Shannon entropy and f-divergences, which disregard structural similarities. By integrating the underlying geometry of the support set, the authors define a structure-aware entropy and derive corresponding Bregman divergences that retain desirable properties of the Kullback–Leibler divergence and Shannon entropy while embedding pairwise similarities directly into the divergence formulation. The approach successfully uncovers structural patterns missed by conventional methods in synthetic clustering tasks, achieves computational efficiency several orders of magnitude higher than optimal transport, and reproduces and extends established findings in applications to economic geography and ecology.