GPU-Accelerated Bregman Douglas-Rachford Splitting for Discrete Optimal Transport

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of low computational efficiency and poor numerical stability in discrete optimal transport across diverse input formats, including cost matrices, point clouds, and meshes. To this end, it proposes a hardware-aware Bregman Douglas-Rachford splitting (BDRS) algorithm that achieves mathematically equivalent and efficient solutions for all three formats by optimizing iterative representations and leveraging GPU parallel computation. Furthermore, this work establishes the first unified evaluation benchmark across solvers. Experimental results demonstrate that the proposed method attains state-of-the-art performance across all input modalities, significantly outperforming existing GPU-based baseline solvers.
📝 Abstract
We present GPU-accelerated Bregman Douglas--Rachford splitting algorithm (BDRS) for discrete optimal transport problem in three input formats: an explicit cost matrix, a point cloud with a ground cost between them, and a separable cost on a regular grid. For each input format, we propose hardware-aware designs of mathematically equivalent representations for the BDRS iterations to enhance numerical stability and empirical runtime. We benchmark the three proposed implementations against eight GPU baseline solvers from the literature on the same device. We demonstrate that our implementations of BDRS achieve state-of-the-art performance on their respective input formats. To the best of our knowledge, this is the first cross-solver study of GPU DOT solvers with a unified measure of optimality.
Problem

Research questions and friction points this paper is trying to address.

Discrete Optimal Transport
GPU Acceleration
Bregman Douglas-Rachford Splitting
Numerical Stability
Innovation

Methods, ideas, or system contributions that make the work stand out.

GPU acceleration
Bregman Douglas-Rachford splitting
discrete optimal transport
hardware-aware design
cross-solver benchmark
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Y
Yifan Xu
Department of Computational Applied Mathematics and Operations Research, Rice University, Houston, TX, USA
Shiqian Ma
Shiqian Ma
Rice University
OptimizationMachine LearningLLM