🤖 AI Summary
This work addresses the joint optimization challenge of efficient beamforming and adversarial user detection in dynamic wireless environments by proposing a Bayesian-guided collaborative reinforcement learning framework—the first to integrate such a mechanism into secure beamforming. Leveraging the 3GPP-standardized channel model, the framework employs Q-learning and SARSA algorithms to dynamically select beam directions, simultaneously ensuring communication capacity and enabling attacker identification. Experimental results demonstrate that the proposed approach significantly outperforms random and exhaustive-search baselines in terms of aggregate channel capacity, detection accuracy, and system stability. Among the two algorithms, Q-learning achieves the best trade-off between precision and computational efficiency.
📝 Abstract
In next-generation wireless networks, communication systems are expected to go beyond simple data transmission and simultaneously provide high data rates, efficiency, and security. This requirement has motivated the extensive adoption of machine learning methods to develop intelligent and real-time network management frameworks, enabling the system to continuously monitor and react to channel variations and user behavior while maintaining efficient information delivery. In this context, the integration of machine learning with beamforming enables adaptive and data-driven beam direction selection, improving both the efficiency and security of wireless links. In this work, a 3GPP-based system model is first implemented under a no-attacker scenario, and an exhaustive search is employed as a reference to identify the best beamforming configurations. The proposed framework is then evaluated in the presence of an attacker and under different network scalability conditions. We demonstrate that the reinforcement learning-based approaches, namely Q-learning and SARSA (State-Action-Reward-State-Action), consistently outperform random selection in terms of total channel capacity, attacker detection accuracy, and performance stability. Among the evaluated reinforcement learning methods, Q-learning achieves the best overall trade-off between detection accuracy and computational efficiency. Our results indicate that the proposed framework provides a stable, scalable, and effective solution for joint beamforming and security-aware decision-making in dynamic and adversarial wireless environments.