π€ AI Summary
This study addresses the computational intractability of joint optimization in cell-free massive MIMO systems with movable antennas, caused by their non-convex constraints. To overcome this challenge, we propose GLIIR-HAPPO, a heterogeneous multi-agent reinforcement learning framework that decomposes the coupled optimization problem into coordinated subproblems. The framework achieves efficient cross-role collaboration and policy synchronization through a dynamic interaction graph critic combined with a role-conditioned federated distillation architecture. Furthermore, solution quality is ensured via a penalty-augmented reward structure integrated with a dedicated geometric solver. Experimental results demonstrate that the proposed method significantly enhances the system sum-rate, approaching the performance of centralized schemes while substantially reducing communication overhead. Additionally, the framework provides theoretical convergence guarantees.
π Abstract
The inherent non-convex minimum-separation constraints introduced by movable antennas present a formidable challenge to the joint optimization of antenna positions and transmission strategies, rendering conventional methods computationally infeasible, particularly in large-scale cell-free massive multiple-input multiple-output (MIMO). In this work, we propose the graph-based learning individual intrinsic reward heterogeneous-agent proximal policy optimization (GLIIR-HAPPO) algorithm, a novel heterogeneous multi-agent reinforcement learning (MARL) framework that fundamentally overcomes this impasse by systematically decomposing the original coupled optimization into coordinated subproblems. To ensure tractability, we embed the non-convex geometric constraints into a penalty-augmented reward structure and develop a specialized geometric solver that enables the positioning agents to efficiently navigate the high-dimensional action space. Specifically, we propose an architecture featuring a dynamic-interaction graph critic for adaptive cross-role coordination, together with role-conditioned federated distillation that synchronizes same-role policies through compact actor-output statistics. Beyond architectural design, we establish a rigorous theoretical analysis that derives monotonic performance improvement bounds and establishes convergence guarantees for the proposed bi-level optimization. Numerical simulations demonstrate that our framework yields significant sum-rate improvements over state-of-the-art MARL schemes. Moreover, the performance of our advanced architecture closely approaches its fully centralized counterpart, while drastically reducing communication overhead.