🤖 AI Summary
This study addresses the challenging deployment of multi-UAV networks in threat environments, where both security and global energy efficiency must be jointly guaranteed. To this end, we propose the TAKM-MATD3 framework. Methodologically, it employs threat-aware K-means clustering with optimal matching to achieve secure UAV grouping, and leverages Multi-Agent Twin Delayed Deep Deterministic Policy Gradient (MATD3) to dynamically optimize trajectories and power allocation. The proposed approach attains performance comparable to global optimization at minimal computational complexity while exhibiting strong generalization capabilities. Experimental results demonstrate that the framework significantly enhances system energy efficiency and convergence speed under a strict zero-security-violation guarantee. Furthermore, its online deployment complexity is substantially lower than that of conventional heuristic methods, making it highly practical for real-time applications.
📝 Abstract
Ensuring operational safety in threat-prone environments remains a critical challenge for multi-UAV networks serving as aerial base stations. This paper proposes an efficient framework to maximize global energy efficiency (EE) while promoting safe operation through threat-aware clustering and reward-based safety enforcement. The proposed framework is executed in three steps. First, a threat-aware K-means (TAKM) algorithm determines the minimum required UAVs and computes safe initial placements. Second, an optimal matching stage assigns physical UAVs to these centroids to minimize energy expenditure. Third, a threat-aware multi-agent twin delayed deep deterministic policy gradient (MATD3) algorithm dynamically optimizes trajectories, power, and user associations. Simulation results show that the proposed framework achieves zero observed safety violations in the considered scenarios while achieving superior EE and faster convergence than other learning methods and non-clustering baselines. Compared to heuristic optimization, the proposed framework outperforms the greedy particle swarm optimization (GPSO) and achieves performance comparable to that of the optimized PSO (OPSO), while incurring significantly lower online deployment computational complexity. Furthermore, the proposed framework demonstrates effective generalization to unseen user distributions, large UAV fleets, and different threat geometries, while maintaining zero safety violations.