Heterogeneous Multi-Agent Reinforcement Learning for Radio Resource Management under Coupled Finite-Horizon Constraints

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the joint optimization of user association, scheduling, base station activation, and handover control in dense wireless networks under finite-time energy and handover constraints, aiming to balance system throughput and proportional fairness. The authors propose HeLyMARL, a novel framework that integrates Lyapunov virtual queues into heterogeneous multi-agent reinforcement learning for the first time. By leveraging drift-plus-penalty decomposition, it internalizes coupled long-term constraints into per-slot rewards, thereby transforming the original constrained problem into an unconstrained learning task. Unlike conventional approaches that enforce constraints only across episodes, HeLyMARL guarantees strict budget adherence over any partial time horizon. Experiments demonstrate that HeLyMARL is the only method capable of maintaining uninterrupted service throughout while simultaneously achieving high throughput and fairness, significantly outperforming existing MARL, Lyapunov-based control, and constrained MARL approaches.
📝 Abstract
Maximizing throughput under proportional fairness in dense wireless networks requires jointly managing user association, scheduling, base station (BS) activation, and handover control under hard finite-horizon energy and handover budgets, which induces a fundamental tension between BS-side energy management and user-side handover regulation. While multi-agent reinforcement learning (MARL) is a natural framework for such distributed sequential control, its application here faces two difficulties: finite-horizon budget constraints cannot be evaluated at each time slot, and the nonlinear proportional fairness utility admits no principled per-slot decomposition. We propose HeLyMARL, a Lyapunov-embedded heterogeneous MARL framework that resolves both via drift-plus-penalty decomposition with virtual queues. The energy and handover constraint pressures are internalized directly into a unified per-slot reward, converting the constrained finite-horizon problem into an unconstrained MARL problem. Comparison against two Lagrangian-based alternatives reveals a timescale separation: Lagrangian relaxation regulates constraints only across training episodes, whereas the virtual queues of HeLyMARL bound cumulative budget consumption at every partial horizon within an episode, a pacing guarantee beyond the reach of greedy Lyapunov-based control. Simulations show that HeLyMARL is the only method that sustains the throughput-fairness balance together with uninterrupted service throughout the horizon, outperforming conventional MARL, Lyapunov-based, and constrained MARL benchmarks without premature budget exhaustion.
Problem

Research questions and friction points this paper is trying to address.

Radio Resource Management
Proportional Fairness
Finite-Horizon Constraints
Multi-Agent Reinforcement Learning
Handover Control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Heterogeneous MARL
Lyapunov optimization
Virtual queues
Finite-horizon constraints
Proportional fairness
🔎 Similar Papers
No similar papers found.