Score
Designs, implements, or analyzes differentiable entropic optimal-transport solvers (Sinkhorn iterations) and their integration as differentiable layers so that the transport plan and related objectives are computed within a gradient-based training pipeline. Uses these differentiable Sinkhorn components to backpropagate through transport computations to learn cost functions or model parameters and to align or match distributions via learned transport.
Optimal transport (OT) faces critical scalability, robustness, and ethical challenges in large-scale data settings. Method: This work systematically surveys OT’s evolution—from the Monge–Kantorovich foundation to modern computational approaches, including Sinkhorn iterations, primal-dual optimization, dimensionality reduction, and problem reduction techniques—and innovatively unifies scalability enhancements with emerging variants, notably Optimal Transport Warping (OTW). Contribution/Results: We present the first theoretically rigorous yet broadly applicable OT framework, achieving cross-domain practicality without sacrificing mathematical soundness. Empirical evaluation demonstrates that OTW significantly outperforms Dynamic Time Warping (DTW) in temporal alignment tasks. Our comprehensive OT algorithmic landscape establishes a new paradigm for high-dimensional distribution comparison—characterized by efficiency, numerical stability, and interpretability. Furthermore, the work proactively identifies key open challenges, including robust statistical modeling under distributional shift and fairness-aware constraints in OT-based inference.
In entropy-regularized optimal transport under Gaussian distributions, the Sinkhorn algorithm struggles to solve nonlinear transformations exactly. Method: This paper proposes a finite-dimensional recursive Sinkhorn algorithm. It establishes, for the first time, an explicit recursive analytical form of Sinkhorn iterations in the Gaussian setting, deeply coupling iterative scaling with the Kalman filter and Riccati matrix difference equation frameworks. Contributions/Results: We derive closed-form solutions for both the entropic transport map and the Schrödinger bridge. Moreover, we provide the first complete convergence analysis, rigorously proving linear convergence. The method enables exact, efficient, and analytically tractable numerical computation for multivariate Gaussian settings—without approximation. By unifying optimal transport, Schrödinger bridges, and filtering theory, it offers a novel tool for probabilistic modeling and dynamic inference.
To address the high computational cost and curse-of-dimensionality challenges in sampling optimal transport (OT) couplings for large-scale, high-dimensional data, this paper proposes an efficient learning framework based on score-based generative models. Specifically, conditioned on source samples, it iteratively generates target samples following the Sinkhorn-regularized OT coupling via Langevin dynamics. Crucially, it jointly parameterizes the score function and Sinkhorn potential functions—enabling, for the first time, end-to-end co-learning of score-based generation and OT coupling. We theoretically establish the convergence of gradient descent on the network parameters under mild assumptions. Experiments demonstrate that our method significantly improves both accuracy and speed of coupling estimation across diverse large-scale OT tasks, while maintaining scalability and practical applicability.
This work investigates the continuous limit behavior of Sinkhorn iterations (i.e., the iterative proportional fitting procedure) in the 2-Wasserstein space as the regularization parameter ε → 0 and the number of iterations scales as 1/ε. Methodologically, we introduce the novel concept of Wasserstein mirror gradient flow, integrating relative entropy gradient analysis, asymptotic expansion of Monge–Ampère-type PDEs, and McKean–Vlasov diffusion construction. We rigorously establish existence and uniqueness of the limiting curve, proving it is absolutely continuous in Wasserstein space; its velocity field admits an explicit characterization via metric derivatives and is exactly reproduced by a first-order stochastic differential equation. This yields the first rigorous connection between Sinkhorn iteration and parabolic optimal transport dynamics, leading to an exponential convergence criterion and revealing Sinkhorn’s intrinsic nature as an implicit gradient flow on the Wasserstein manifold.
Traditional reduced-order models (ROMs) struggle to accurately capture the geometric structure of high-dimensional solution manifolds due to slow decay of the Kolmogorov *n*-width. To address this, we propose a novel nonlinear ROM framework integrating optimal transport (OT) with deep learning. Our method introduces a Wasserstein kernel by embedding the Wasserstein distance into the kernel function of kernel proper orthogonal decomposition (kPOD), and employs the Sinkhorn divergence as the loss function for neural network training—enabling OT-aware nonlinear dimensionality reduction. This approach significantly enhances geometric fidelity to slowly decaying manifolds. Compared to conventional ROMs, it achieves superior accuracy, improved training stability, enhanced robustness to noise, faster convergence, and higher computational efficiency.
This paper investigates the bidirectional causal optimal transport (OT) problem with adapted structural coupling. To exploit its dynamic programming structure, we propose the first fitted value iteration (FVI) framework, employing deep neural networks to approximate the value function. Theoretically, under assumptions of concentrability and approximation completeness, we derive a sample complexity upper bound based on local Rademacher complexity and verify that suitably structured neural networks satisfy the required conditions. Experimentally, our method significantly outperforms linear programming and adapted Sinkhorn algorithms in computational efficiency as the time horizon increases, while maintaining controllable accuracy; it further exhibits strong scalability and practical feasibility. The core contribution lies in systematically introducing FVI to bidirectional causal OT—thereby establishing the first model-free approximate solution framework for this problem and filling a critical theoretical and methodological gap in the literature.
本文提出SinkSLOT方法,通过稀疏提升的运输计划解决大规模数据集上熵最优传输计算效率低和独立耦合问题。
研究通过固定支持图控制稀疏Sinkhorn层中梯度传播的问题,利用固定支持计算和Dobrushin收缩等方法分析并设定了设计可微传输层的支持图的数学标准。
This work addresses the limitations of traditional Bayesian inference, which relies on exact likelihoods and suffers when the likelihood is misspecified, intractable, or misaligned with the target discrepancy. The authors propose the first integration of Sinkhorn divergence as a generalized Bayesian loss within Hamiltonian Monte Carlo (HMC) and its adaptive variant NUTS, accommodating both balanced and unbalanced optimal transport settings. They further incorporate common random numbers to handle stochastic simulators and introduce a heuristic for hyperparameter selection that preserves gradient consistency with Hamiltonian dynamics. Empirical evaluations on Gaussian models, noisy spiral manifolds, pulse misalignment, and CIFAR-10 image patch alignment demonstrate the method’s effectiveness, robustness, and the nuanced differences between transport mechanisms.
This work addresses the limitation of traditional causal inference methods—such as average treatment effects—which capture only local differences in outcome distributions and thus fail to fully characterize the treatment’s impact on the entire distribution. The authors propose the Sinkhorn Treatment Effect, which for the first time integrates entropy-regularized optimal transport into causal inference. By constructing a smooth transformation of counterfactual mean embeddings, they derive a differentiable functional representation of distributional treatment effects. Building on this framework, they develop a debiased estimator with asymptotic efficiency and a multi-regularization-parameter aggregation test. Both theoretical analysis and empirical experiments demonstrate that the proposed approach substantially enhances the identification and detection of distributional causal effects on synthetic and image data.
This work addresses the lack of a unified theoretical foundation for parameter identifiability, estimation consistency, and algorithmic convergence in feature-parameterized inverse optimal transport (IOT). It proposes an analytical framework based on Sinkhorn linearization and spectral surrogates, which—through the introduction of spectral sandwich inequalities, restricted Hessian analysis, and ℓ₁ regularization—establishes, for the first time, four core theoretical guarantees for IOT: global parameter identifiability, exact support recovery, a Lipschitz continuous inverse mapping, and monotonic convergence of gradient descent. The framework combines geometric transparency with spectral precision and further quantifies the Hölder stability of the projection map under model misspecification. Numerical experiments corroborate the theoretical findings.