π€ AI Summary
This work addresses the instability in multi-agent reinforcement learning caused by dynamic reward weights generated by large language models (LLMs), which violate the stationarity assumption of potential-based reward shaping (PBRS), contaminate experience replay, and destabilize training. The study is the first to identify three distinct failure modes arising from LLM-induced reward dynamics that disrupt mechanistic dependencies in the learning process. To mitigate these issues, the authors propose two stabilization strategies: training-phase-dependent weight freezing and exponential moving average (EMA) smoothing of reward signals. Integrated with QMIX and VDN architectures, the approach achieves success rates of 86.7%, 95.9%, and 99.9% on the Simple Spread, Level-Based Foraging, and SMAC 3m benchmarks, respectively, substantially outperforming baseline methods. These results establish reward signal stationarity as a critical design constraint for LLM-augmented multi-agent systems.
π Abstract
Large Language Models (LLMs) offer a natural interface for translating human objectives into reward signals for cooperative multi-agent reinforcement learning (MARL), yet the training-time dynamics of this integration remain poorly understood. We show that dynamically updating LLM-generated reward weights during off-policy MARL violates the stationarity assumption of Potential-Based Reward Shaping (PBRS) and contaminates the experience replay buffer, whose stored transitions carry reward labels computed under stale shaping weights. We characterise the result as a regime-dependent failure whose severity depends on how competent the unshaped baseline already is. To control it we propose two stabilisation strategies: a Phase-Based Freeze Schedule that enforces strict stationarity within training phases, and Exponential Moving Average (EMA) smoothing that bounds per-episode weight drift. We evaluate across three cooperative environments and five random seeds with QMIX, complemented by an exploratory VDN extension, yielding a three-regime taxonomy. In the augmentative regime (Simple Spread), where the baseline is functional (74.4 %), EMA significantly improves success to 86.7 % ($+12.3$ pp, $p<0.01$) while naive dynamic updates collapse it to 15.2 %. In the essential regime (Level-Based Foraging), where the baseline is broken (0.1 %), any shaping unlocks the task (95.9 % under EMA). In the supplementary regime (SMAC 3m), where the baseline is near-saturated (98.8 %), stabilised shaping preserves performance (99.9 %) while unstabilised shaping adds variance without gain. These findings establish reward-signal stationarity as a necessary design constraint and indicate that regime placement is a practical predictor of whether dynamic LLM shaping helps or harms.