differentiable sinkhorn

Designs, implements, or analyzes differentiable entropic optimal-transport solvers (Sinkhorn iterations) and their integration as differentiable layers so that the transport plan and related objectives are computed within a gradient-based training pipeline. Uses these differentiable Sinkhorn components to backpropagate through transport computations to learn cost functions or model parameters and to align or match distributions via learned transport.

differentiablesinkhorn

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.39
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Gaussian entropic optimal transport: Schr""odinger bridges and the Sinkhorn algorithm

Dec 24, 2024
OD
O. Deniz Akyildiz
🏛️ Imperial College London | Centre de Recherche Inria Bordeaux Sud-Ouest | Universidad Carlos III de Madrid

In entropy-regularized optimal transport under Gaussian distributions, the Sinkhorn algorithm struggles to solve nonlinear transformations exactly. Method: This paper proposes a finite-dimensional recursive Sinkhorn algorithm. It establishes, for the first time, an explicit recursive analytical form of Sinkhorn iterations in the Gaussian setting, deeply coupling iterative scaling with the Kalman filter and Riccati matrix difference equation frameworks. Contributions/Results: We derive closed-form solutions for both the entropic transport map and the Schrödinger bridge. Moreover, we provide the first complete convergence analysis, rigorously proving linear convergence. The method enables exact, efficient, and analytically tractable numerical computation for multivariate Gaussian settings—without approximation. By unifying optimal transport, Schrödinger bridges, and filtering theory, it offers a novel tool for probabilistic modeling and dynamic inference.

Gaussian EntropyOptimal TransportSinkhorn Algorithm

Score-based Generative Neural Networks for Large-Scale Optimal Transport

Oct 07, 2021
MD
Max Daniels
🏛️ Northeastern University | Brandeis University

To address the high computational cost and curse-of-dimensionality challenges in sampling optimal transport (OT) couplings for large-scale, high-dimensional data, this paper proposes an efficient learning framework based on score-based generative models. Specifically, conditioned on source samples, it iteratively generates target samples following the Sinkhorn-regularized OT coupling via Langevin dynamics. Crucially, it jointly parameterizes the score function and Sinkhorn potential functions—enabling, for the first time, end-to-end co-learning of score-based generation and OT coupling. We theoretically establish the convergence of gradient descent on the network parameters under mild assumptions. Experiments demonstrate that our method significantly improves both accuracy and speed of coupling estimation across diverse large-scale OT tasks, while maintaining scalability and practical applicability.

Learning Sinkhorn coupling via score-based networksSampling optimal transport coupling between distributionsSolving high-dimensional transport without linear programming

This work investigates the continuous limit behavior of Sinkhorn iterations (i.e., the iterative proportional fitting procedure) in the 2-Wasserstein space as the regularization parameter ε → 0 and the number of iterations scales as 1/ε. Methodologically, we introduce the novel concept of Wasserstein mirror gradient flow, integrating relative entropy gradient analysis, asymptotic expansion of Monge–Ampère-type PDEs, and McKean–Vlasov diffusion construction. We rigorously establish existence and uniqueness of the limiting curve, proving it is absolutely continuous in Wasserstein space; its velocity field admits an explicit characterization via metric derivatives and is exactly reproduced by a first-order stochastic differential equation. This yields the first rigorous connection between Sinkhorn iteration and parabolic optimal transport dynamics, leading to an exponential convergence criterion and revealing Sinkhorn’s intrinsic nature as an implicit gradient flow on the Wasserstein manifold.

Analyzing convergence of Sinkhorn algorithm iterations to Wasserstein flowCharacterizing exponential convergence conditions for limiting Sinkhorn flowEstablishing connection between Sinkhorn algorithm and parabolic Monge-Ampère PDE

Optimal Transport-inspired Deep Learning Framework for Slow-Decaying Problems: Exploiting Sinkhorn Loss and Wasserstein Kernel

Aug 26, 2023
MK
M. Khamlich
🏛️ SISSA | École Polytechnique Fédérale de Lausanne

Traditional reduced-order models (ROMs) struggle to accurately capture the geometric structure of high-dimensional solution manifolds due to slow decay of the Kolmogorov *n*-width. To address this, we propose a novel nonlinear ROM framework integrating optimal transport (OT) with deep learning. Our method introduces a Wasserstein kernel by embedding the Wasserstein distance into the kernel function of kernel proper orthogonal decomposition (kPOD), and employs the Sinkhorn divergence as the loss function for neural network training—enabling OT-aware nonlinear dimensionality reduction. This approach significantly enhances geometric fidelity to slowly decaying manifolds. Compared to conventional ROMs, it achieves superior accuracy, improved training stability, enhanced robustness to noise, faster convergence, and higher computational efficiency.

High-Dimensional ProblemsKolmogorov n-widthReduced Order Models

Fitted Value Iteration Methods for Bicausal Optimal Transport

Jun 22, 2023
EB
Erhan Bayraktar
🏛️ University of Michigan

This paper investigates the bidirectional causal optimal transport (OT) problem with adapted structural coupling. To exploit its dynamic programming structure, we propose the first fitted value iteration (FVI) framework, employing deep neural networks to approximate the value function. Theoretically, under assumptions of concentrability and approximation completeness, we derive a sample complexity upper bound based on local Rademacher complexity and verify that suitably structured neural networks satisfy the required conditions. Experimentally, our method significantly outperforms linear programming and adapted Sinkhorn algorithms in computational efficiency as the time horizon increases, while maintaining controllable accuracy; it further exhibits strong scalability and practical feasibility. The core contribution lies in systematically introducing FVI to bidirectional causal OT—thereby establishing the first model-free approximate solution framework for this problem and filling a critical theoretical and methodological gap in the literature.

Computing bicausal optimal transport with adapted coupling structuresDeveloping scalable methods outperforming linear programming approachesEstablishing sample complexity using Rademacher complexity theory

Latest Papers

What's happening recently
View more

研究通过固定支持图控制稀疏Sinkhorn层中梯度传播的问题,利用固定支持计算和Dobrushin收缩等方法分析并设定了设计可微传输层的支持图的数学标准。

Gradient PropagationScaling IterationsSparse Sinkhorn Layers

This work addresses the limitations of traditional Bayesian inference, which relies on exact likelihoods and suffers when the likelihood is misspecified, intractable, or misaligned with the target discrepancy. The authors propose the first integration of Sinkhorn divergence as a generalized Bayesian loss within Hamiltonian Monte Carlo (HMC) and its adaptive variant NUTS, accommodating both balanced and unbalanced optimal transport settings. They further incorporate common random numbers to handle stochastic simulators and introduce a heuristic for hyperparameter selection that preserves gradient consistency with Hamiltonian dynamics. Empirical evaluations on Gaussian models, noisy spiral manifolds, pulse misalignment, and CIFAR-10 image patch alignment demonstrate the method’s effectiveness, robustness, and the nuanced differences between transport mechanisms.

entropic optimal transportGeneralized Bayeslikelihood misspecification

This work addresses the limitation of traditional causal inference methods—such as average treatment effects—which capture only local differences in outcome distributions and thus fail to fully characterize the treatment’s impact on the entire distribution. The authors propose the Sinkhorn Treatment Effect, which for the first time integrates entropy-regularized optimal transport into causal inference. By constructing a smooth transformation of counterfactual mean embeddings, they derive a differentiable functional representation of distributional treatment effects. Building on this framework, they develop a debiased estimator with asymptotic efficiency and a multi-regularization-parameter aggregation test. Both theoretical analysis and empirical experiments demonstrate that the proposed approach substantially enhances the identification and detection of distributional causal effects on synthetic and image data.

causal inferencecounterfactual distributionsdistributional divergence

This work addresses the lack of a unified theoretical foundation for parameter identifiability, estimation consistency, and algorithmic convergence in feature-parameterized inverse optimal transport (IOT). It proposes an analytical framework based on Sinkhorn linearization and spectral surrogates, which—through the introduction of spectral sandwich inequalities, restricted Hessian analysis, and ℓ₁ regularization—establishes, for the first time, four core theoretical guarantees for IOT: global parameter identifiability, exact support recovery, a Lipschitz continuous inverse mapping, and monotonic convergence of gradient descent. The framework combines geometric transparency with spectral precision and further quantifies the Hölder stability of the projection map under model misspecification. Numerical experiments corroborate the theoretical findings.

feature-parameterized costinverse optimal transportmodel misspecification

Hot Scholars

AH

Abhishek Halder

Associate Professor, Iowa State University
Systems and ControlOptimizationProbabilityMachine Learning
PH

Paul Hand

Associate Professor of Mathematics and Computer Science, Northeastern University
Signal recoveryalgorithmsmachine learningmachine vision
YB

Yikun Bai

Postdoc, Computer Science department, Vanderbilt University
optimal transportmachine learningInformation theory
SK

Soheil Kolouri

Computer Science, Vanderbilt University, Nashville, TN
Machine LearningOptimal TransportComputer Vision
LZ

Liang Zheng

ANU | Canva
Computer VisionRepresentation LearningObject Re-identificationGenerative AI