Who Bears the Burden? Learning Responsibility for Shared Constraints in Multi-Agent Reinforcement Learning

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the imbalanced penalty allocation caused by a uniform Lagrange multiplier under shared cost constraints in multi-agent systems. To this end, we propose LiRA, a method that learns each agent's responsibility share through social welfare optimization to redistribute the penalty burden of shared constraints, without altering the original rewards or constraints. Furthermore, we derive a welfare gradient that accounts for data distribution shifts, thereby inducing a smooth family of normalized generalized Nash equilibria. Evaluations on benchmarks such as CityLearn demonstrate that LiRA improves average social welfare by up to 29%, while effectively reducing violation rates and enhancing budget utilization.
📝 Abstract
When multiple agents share a cost budget, a common Lagrange multiplier can enforce the aggregate constraint but does not determine how its penalty should be allocated across agents. Uniform penalties ignore heterogeneity in the rewards agents sacrifice, while agent-specific multipliers may still rely on the same aggregate cost signal. We introduce Lagrangian Responsibility Allocation (LiRA), which learns each agent's share of a common multiplier by optimizing social welfare over a finite training horizon. The multiplier enforces the aggregate budget, while responsibility shares redistribute its influence without modifying the original rewards or constraints. For convex games under standard regularity conditions, varying these shares induces a smooth family of normalized generalized Nash equilibria in which active constraints remain at their budgets while welfare varies. To optimize responsibility before convergence, we derive a welfare gradient that accounts for both learning updates and the induced change in data distribution. Across CityLearn, MABIM, Harvest, and MetaDrive, spanning 3 to 400 agents, LiRA improves average social welfare by up to 29% over uniform and agent-specific multiplier baselines. Grid and driving costs remain within budget, inventory violations decrease, and Harvest makes more effective use of available budget.
Problem

Research questions and friction points this paper is trying to address.

Multi-Agent Reinforcement Learning
Shared Constraints
Responsibility Allocation
Lagrange Multiplier
Social Welfare
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Agent Reinforcement Learning
Lagrangian Responsibility Allocation
Generalized Nash Equilibrium
Social Welfare Optimization
Shared Constraints
🔎 Similar Papers
No similar papers found.
X
Xiaoyang Cao
Massachusetts Institute of Technology
Jingqi Li
Jingqi Li
Postdoc at UT Austin
Dynamic gamescontrol theorydeep RL
Z
Zhe Fu
Stanford University
A
Alexandre M. Bayen
University of California, Berkeley