🤖 AI Summary
This study addresses the joint optimization of access point operating modes, reconfigurable intelligent surface (RIS) phase shifts, and power allocation in a cell-free massive MIMO-SWIPT system enhanced by stacked RISs, under practical constraints including nonlinear energy harvesting, pilot contamination, channel estimation errors, and reliance solely on long-term statistical channel state information. The work aims to simultaneously maximize harvested energy while meeting spectral efficiency requirements. For the first time, stacked RISs are introduced into this context, and a centralized training–decentralized execution (CTDE) deep reinforcement learning framework is proposed. By modeling the environment as a Markov decision process and employing a normalized joint reward function, the method efficiently tackles the high-dimensional mixed-integer nonconvex optimization problem. Experimental results demonstrate that the proposed approach significantly outperforms conventional schemes under dual constraints and achieves performance close to the convex optimization theoretical upper bound, confirming its effectiveness and robustness.
📝 Abstract
This study explores a next-generation multiple access (NGMA) framework for cell-free massive MIMO (CF-mMIMO) systems enhanced by stacked intelligent metasurfaces (SIMs), aiming to improve simultaneous wireless information and power transfer (SWIPT) performance. A fundamental challenge lies in optimally selecting the operating modes of access points (APs) to jointly maximize the received energy and satisfy spectral efficiency (SE) quality-of-service constraints. Practical system impairments, including a non-linear harvested energy model, pilot contamination (PC), channel estimation errors, and reliance on long-term statistical channel state information (CSI), are considered. We derive closed-form expressions for both the achievable SE and the average sum harvested energy (sum-HE). A mixed-integer non-convex optimization problem is formulated to jointly optimize the SIM phase shifts, APs mode selection, and power allocation to maximize average sum-HE under SE and average harvested energy constraints. To solve this problem, we propose a centralized training, decentralized execution (CTDE) framework based on deep reinforcement learning (DRL), which efficiently handles high-dimensional decision spaces. A Markovian environment and a normalized joint reward function are introduced to enhance the training stability across on-policy and off-policy DRL algorithms. Additionally, we provide a two-phase convex-based solution as a theoretical robust performance. Numerical results demonstrate that the proposed DRL-based CTDE framework achieves SWIPT performance comparable to convexification-based solution, while significantly outperforming baselines.