Interactive Distributionally Robust Multi-Agent Learning with General Function Approximation

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that model misspecification and strategic interactions readily amplify transition uncertainty in multi-agent reinforcement learning. To tackle this, we propose the RoMEX-φ framework, which leverages φ-divergence uncertainty sets and their dual representations to achieve distributionally robust online learning under general function approximation. The method integrates equilibrium exploration with dual fitting, introducing a robust multi-agent decoupling coefficient alongside a centered empirical robust difference estimation technique. By replacing state-space cardinality with function class complexity, we establish sublinear robust regret bounds. Large-scale experiments demonstrate that the proposed framework significantly enhances resilience against transition shifts, outperforming non-robust baselines while closely approximating the performance of exact tabular methods.
📝 Abstract
Model misspecification poses a fundamental challenge in multi-agent reinforcement learning, where transition uncertainty can be amplified by strategic interactions among agents. Distributionally robust Markov games (DRMGs) provide a principled framework for addressing such uncertainty, yet existing methods often rely on restrictive assumptions or scale poorly to large state and joint action spaces. We study online learning in general-sum DRMGs with general function approximation and $\phi$-divergence uncertainty sets. We propose RoMEX-$\phi$, a model-free framework that integrates equilibrium-based exploration with dual fitted learning. Through a functional dual representation of the robust multi-agent Bellman operator, RoMEX-$\phi$ enables tractable worst-case value estimation from nominal interaction data using a centered empirical robust discrepancy. We introduce the robust Multi-Agent Decoupling Coefficient (robust MADC) to characterize the intrinsic exploration complexity arising from strategic interactions and adversarial transition uncertainty. We establish sublinear robust regret guarantees governed by the robust MADC rather than explicitly by the state and joint action space sizes, replacing tabular dependence with intrinsic function-class complexity. Numerical experiments on a scalable general-sum DRMG under total variation uncertainty show that RoMEX-$\phi$ is substantially more resilient to transition shifts than its non-robust counterpart while remaining competitive with an exact tabular robust baseline. Our results provide a scalable framework for distributionally robust multi-agent reinforcement learning with general function approximation.
Problem

Research questions and friction points this paper is trying to address.

multi-agent reinforcement learning
distributionally robust Markov games
model misspecification
general function approximation
transition uncertainty
Innovation

Methods, ideas, or system contributions that make the work stand out.

Distributionally Robust Markov Games
General Function Approximation
Model-Free Learning
Robust Multi-Agent Decoupling Coefficient
Regret Bounds
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
D
Debamita Ghosh
Department of Electrical and Computer Engineering, University of Central Florida, Orlando, FL 32816, USA
G
George K. Atia
Department of Electrical and Computer Engineering, University of Central Florida, Orlando, FL 32816, USA; Department of Computer Science, University of Central Florida, Orlando, FL 32816, USA
Yue Wang
Yue Wang
University of Central Florida
Reinforcement LearningOptimizationgame theory