rl-based beamforming

Designs, trains, and evaluates beamforming policies or precoding controllers using reinforcement learning algorithms that output antenna-array weight vectors or transmit/receive strategies. Implements online or offline RL training, reward design, and performance analysis to optimize metrics such as SNR, capacity, robustness, or secrecy under uncertain, time-varying channels and adversarial conditions.

rl-basedbeamforming

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.3
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the joint optimization challenge of efficient beamforming and adversarial user detection in dynamic wireless environments by proposing a Bayesian-guided collaborative reinforcement learning framework—the first to integrate such a mechanism into secure beamforming. Leveraging the 3GPP-standardized channel model, the framework employs Q-learning and SARSA algorithms to dynamically select beam directions, simultaneously ensuring communication capacity and enabling attacker identification. Experimental results demonstrate that the proposed approach significantly outperforms random and exhaustive-search baselines in terms of aggregate channel capacity, detection accuracy, and system stability. Among the two algorithms, Q-learning achieves the best trade-off between precision and computational efficiency.

adversarial user detectionbeamformingchannel capacity

This work addresses the limitations of conventional millimeter-wave and terahertz systems that rely on predefined beam codebooks, which suffer from degraded performance and poor robustness under non-ideal conditions such as non-line-of-sight propagation, hardware impairments, and feedback noise. To overcome these challenges, the paper proposes a multi-agent reinforcement learning framework that learns robust beam codebooks directly from environmental feedback without requiring prior channel state information. It presents the first systematic evaluation of stochastic policies for this task, implementing and comparing three off-policy algorithms—DDPG, TD3, and SAC—and demonstrates that the stochastic-policy-based SAC significantly enhances stability and adaptability. Simulations show that SAC maintains high beamforming gain and training robustness even in harsh scenarios involving strong hardware impairments, non-line-of-sight conditions, and high feedback noise, thereby surpassing the limitations of traditional deterministic approaches.

beam codebookshardware impairmentsmmWave/THz systems

Radio resource management (RRM) in intelligent wireless networks faces challenges in conducting online environment interaction, while existing reinforcement learning (RL) methods inadequately model environmental uncertainty and decision risk. Method: This paper proposes an offline distributional RL framework—integrating offline RL with distributional RL (e.g., IQN or QR-DQN)—that trains exclusively on static historical datasets and explicitly models the return distribution to capture channel stochasticity and operational risk, eliminating reliance on online interaction entirely. Contribution/Results: Evaluated under realistic channel models, the method significantly outperforms conventional heuristics and online RL baselines; it achieves a 10% performance gain over the best online RL approach. To our knowledge, it is the first RRM solution that surpasses state-of-the-art online RL under strictly offline settings.

Intelligent Wireless NetworksRadio Resource ManagementReinforcement Learning

Offline and Distributional Reinforcement Learning for Wireless Communications

Apr 04, 2025
EE
Eslam Eldeeb
🏛️ University of Oulu

In 6G networks, high mobility, strong uncertainty, and stringent real-time requirements render conventional online reinforcement learning (RL) impractical due to its reliance on risky, environment-dependent interactions. Method: This paper proposes an intelligent control framework integrating offline RL with distributional RL, introducing Conservative Quantile Regression (CQR)—a novel risk-sensitive, sample-efficient algorithm—for joint UAV trajectory optimization and wireless resource management. By eliminating online interaction, the framework enhances privacy, robustness, and decision safety while enabling training solely on static datasets. Contribution/Results: Leveraging deterministic policy gradients, our method achieves 37% faster convergence than standard RL baselines; reduces the probability of high-risk actions significantly; and satisfies latency constraints with ≥99.2% reliability—demonstrating superior safety, efficiency, and practicality for 6G network control.

Addressing uncertainties in wireless applications with distributional RLEnhancing scalability and reliability in 6G networks via offline RLOvercoming limitations of online RL in real-time wireless networks

Dynamic reactive jamming attacks—where jammers sense the environment in real time and adaptively select channels and detection thresholds to disrupt communications—pose significant challenges to cognitive radio networks. Method: This paper proposes a hybrid deep reinforcement learning (DRL) framework for anti-jamming cognitive radios, integrating Q-learning (for discrete jamming events) with Deep Q-Networks (to model continuous states such as received power), jointly optimizing transmit power, modulation scheme, and channel selection under unknown channel dynamics and jamming strategies. Contributions/Results: First, it introduces a dual-granularity state representation within a unified DRL architecture—the first of its kind. Second, it designs a multi-objective reward function balancing throughput, energy efficiency, and robustness. Experiments demonstrate rapid convergence under rapidly evolving jamming policies, a 23.6% improvement in long-term throughput, and superior spectral efficiency and anti-jamming robustness compared to baseline methods.

Adapting transmission parameters to counter evolving jamming strategiesMitigating reactive jamming with dynamic channel selectionOptimizing throughput using reinforcement learning without prior knowledge

Latest Papers

What's happening recently
View more

This work addresses physical-layer security in low Earth orbit satellite uplink communications under multiple non-colluding eavesdroppers by proposing a secrecy rate maximization method subject to an average outage constraint. Exploiting the predictability of orbital dynamics, the problem is formulated as a constrained Markov decision process. A primal-dual soft actor-critic reinforcement learning algorithm is developed for beamforming, incorporating a differentiable upper bound on the outage probability as the cost function and employing a multi-head cost critic to enable low-complexity real-time optimization. Theoretical analysis leverages closed-form outage probability expressions under Nakagami-m fading, with constraints handled via Lagrangian relaxation. Experimental results demonstrate that the proposed approach consistently outperforms maximal ratio transmission across diverse eavesdropping configurations and surpasses zero-forcing beamforming in dense eavesdropper scenarios, achieving 93% of the secrecy rate attained by an offline successive convex approximation benchmark with only a single forward computation.

LEO satellite communicationsoutage constraintsphysical-layer security

This study addresses the lack of systematic evaluation of offline reinforcement learning (Offline RL) algorithms under realistic stochastic wireless environments characterized by channel fading, noise, and traffic mobility. For the first time, it presents a comprehensive comparison of Bellman-based Conservative Q-Learning (CQL), the sequence modeling approach Decision Transformer (DT), and their hybrid variant Critic-Guided DT within the open-source communication simulation platform mobile-env. Experimental results demonstrate that CQL exhibits the strongest robustness across various stochastic perturbations, making it a reliable default choice, whereas DT-based methods can outperform traditional Bellman approaches when sufficient high-quality trajectories are available. These findings provide empirical evidence and practical guidance for algorithm selection in AI-driven network control systems such as O-RAN and 6G.

Algorithm SelectionOffline Reinforcement LearningRobustness

This work addresses the challenge of integrating synthetic aperture radar (SAR) sensing with secure communication in wideband systems under emergency or surveillance scenarios where the location of a mobile eavesdropper is unknown. The authors propose a dynamic time-division joint framework that leverages an airborne base station to estimate the ground-based eavesdropper’s position and velocity in real time via cognitive SAR along-track interferometry. These estimates are then used to adaptively optimize beamforming, artificial noise injection, and time-power allocation to maximize the worst-case secrecy rate. For the first time, deep reinforcement learning is introduced to this joint optimization problem, formulated as a Markov decision process, enabling online adaptation and generalization to unknown eavesdropping trajectories. Simulations demonstrate that the proposed method significantly outperforms baseline schemes with equal aperture or random time-slot allocation in terms of secrecy rate and effectively generalizes to unseen eavesdropper motion patterns.

aerial base stationmobile eavesdropperSAR-communication integration

This work proposes a novel offline multi-agent reinforcement learning (MARL) framework that integrates Conservative Q-Learning (CQL) with meta-learning to address the high cost, weak safety guarantees, and poor scalability of conventional online MARL in complex 6G networks. By uniquely combining offline MARL with meta-learning, the approach enhances both the safety of training and the ability to rapidly adapt to dynamic wireless environments. Experimental evaluations in wireless resource management and unmanned aerial vehicle (UAV) network scenarios demonstrate that the proposed method significantly outperforms existing solutions, highlighting the substantial potential and advantages of offline MARL for future 6G communication systems.

6G communicationsnetwork complexityoffline multi-agent reinforcement learning

Hot Scholars

FG

François Grondin

Associate Professor, Université de Sherbrooke
microphone arraydistant speech recognitionrobot auditionsound source localization