π€ AI Summary
This work addresses the energy harvesting efficiency maximization problem for an unmanned aerial vehicle (UAV)-mounted reconfigurable intelligent surface (RIS) communication system, jointly optimizing user transmit power, RIS phase shifts, and time-switching ratios under a nonlinear energy harvesting model and UAV attitude jitter constraints. The resulting optimization problem is highly non-convex and temporally coupled, rendering conventional methods inapplicable. To tackle this challenge, we propose a Smoothed Softmax Double Deep Deterministic Policy Gradient (SS-DDPG) algorithm, incorporating action clipping, entropy regularization, and Softmax-weighted Q-value estimation to enhance policy stability and convergence robustness. Simulation results demonstrate that the proposed algorithm achieves stable convergence across diverse jitter scenarios, attaining an average energy efficiency of 45.07%βclosely approaching the exhaustive-search upper bound of 53.09%βand significantly outperforming existing deep reinforcement learning baselines.
π Abstract
In this letter, we propose an energy-efficient design for an unmanned aerial vehicle (UAV)-mounted reconfigurable intelligent surface (RIS) communication system with nonlinear energy harvesting (EH) and UAV jitter. A joint optimization problem is formulated to maximize the EH efficiency of the UAV-mounted RIS by controlling the user powers, RIS phase shifts, and time-switching factor, subject to quality of service and practical EH constraints. The problem is nonconvex and time-coupled due to UAV angular jitter and nonlinear EH dynamics, making it intractable for conventional optimization methods. To address this, we reformulate the problem as a deep reinforcement learning (DRL) environment and develop a smoothed softmax dual deep deterministic policy gradient algorithm. The proposed method incorporates action clipping, entropy regularization, and softmax-weighted Q-value estimation to improve learning stability and exploration. Simulation results show that the proposed algorithm converges reliably under various UAV jitter levels and achieves an average EH efficiency of 45.07%, approaching the 53.09% upper bound of exhaustive search, and outperforming other DRL baselines.