Optimal Transport Meets Reinforcement Learning: A Survey

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the failure of traditional divergence measures in reinforcement learning (RL) caused by weakly overlapping state distributions by introducing optimal transport (OT) theory. Through a systematic review of OT's application motivations, objective functions, and algorithmic designs in RL, this work proposes a unified framework that integrates OT theory with reinforcement learning. The primary contributions include establishing a comprehensive taxonomy that incorporates practical considerations, providing an in-depth analysis of the core role of OT in policy optimization, and explicitly identifying open challenges such as scalability and theoretical analysis. Ultimately, this research offers clear guidance for future investigations within this interdisciplinary domain.
📝 Abstract
Reinforcement learning (RL) algorithms frequently compare probability distributions, such as state visitation distributions induced by policies and experts, action distributions from learned policies and offline datasets, or transition distributions from learned models and environments. However, commonly used divergences may become ineffective when these distributions overlap weakly, which is frequently encountered in imitation learning, offline RL, and deployment under distribution shift. Optimal transport (OT) offers an alternative by measuring the cost of \emph{moving} probability mass from one distribution to another under a ground cost that encodes task geometry. This survey covers how OT is used inside RL objectives and algorithms. For each method, we identify: the role OT plays, the distributions compared, the OT formulation used, and the treatment of temporal structure. Beyond categorising existing methods, we discuss the motivations behind different OT choices, practical considerations such as cost design and computational challenges, and highlight open problems including scalable trajectory-level transport, principled handling of mass mismatch, and theoretical analysis for OT-regularised RL.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Optimal Transport
Probability Distribution Comparison
Distribution Shift
Imitation Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Optimal Transport
Reinforcement Learning
Distribution Shift
Imitation Learning
Trajectory-level Transport
🔎 Similar Papers
No similar papers found.