wind-aware reinforcement learning

Designs, trains, and evaluates control systems and reinforcement-learning policies that explicitly estimate, condition on, or compensate for wind disturbances so as to reduce trajectory-tracking error and improve robustness across varying wind speeds and turbulence; includes disturbance-aware state representations, disturbance observers, and policy adaptation mechanisms for deployment under changing wind conditions.

wind-awarereinforcementlearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.33
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the significant degradation in trajectory tracking performance of small quadrotors operating in the atmospheric boundary layer under strong turbulent winds. To mitigate this, the authors propose a two-stage learning framework: first, an attention-augmented gated recurrent network leverages onboard kinematic and dynamic data to accurately estimate the local three-dimensional wind field, achieving a horizontal wind speed RMSE of 0.40 m/s and a direction error of 3.2°; second, these wind estimates are integrated into a Proximal Policy Optimization (PPO) reinforcement learning controller to enable wind-aware flight control. This approach represents the first integration of learned wind-field perception with reinforcement learning–based control, reducing trajectory tracking error by 48% compared to a non-wind-aware PD controller across wind speeds of 4–12 m/s, outperforming it in all evaluated scenarios, and maintaining stable flight even under out-of-distribution wind conditions of 13–15 m/s, thereby substantially enhancing wind-robustness.

atmospheric turbulenceflight controlquadrotor

Model-Based Reinforcement Learning for Control of Strongly-Disturbed Unsteady Aerodynamic Flows

Aug 26, 2024
ZL
Zhecheng Liu
🏛️ University of California, Los Angeles | California Institute of Technology

Reinforcement learning (RL) for unsteady aerodynamic flow control under strong disturbances suffers from prohibitively high training costs and poor generalizability to full-scale computational fluid dynamics (CFD) environments. Method: This paper proposes a physics-enhanced model-based RL (MBRL) framework. Its core innovations include: (i) a physics-constrained autoencoder for high-fidelity flow field dimensionality reduction; and (ii) latent-space long-horizon dynamics modeling to improve prediction robustness. Contribution/Results: To our knowledge, this is the first successful application of MBRL to real-world aerodynamic control. In a pitch-controlled airfoil subject to gust disturbances, the learned policy significantly suppresses lift fluctuations. Crucially, the policy transfers seamlessly to full-scale CFD simulations, reducing required training samples by one to two orders of magnitude. The framework establishes a new paradigm for intelligent control of high-dimensional unsteady flows—efficient, physically interpretable, and broadly transferable.

Control of strongly-disturbed unsteady aerodynamic flowsDevelopment of a model-based reinforcement learning approachHigh training cost in model-free reinforcement learning

Linear controllers fail under strong gusts due to nonlinear flow-field interactions. Method: This paper proposes a Transformer-based deep reinforcement learning (DRL) framework for nonlinear aerodynamic lift control using sparse surface pressure feedback. It innovatively integrates expert-policy pretraining—using a linear controller as prior knowledge—with task-level transfer learning (from single- to multi-gust scenarios). The framework identifies that pitch control at the quarter-chord point dominantly modulates added-mass effects, enabling low-energy, high-precision lift regulation. Results: The learned policy significantly outperforms optimal proportional control, with performance gains increasing with gust count. It achieves zero-shot generalization to arbitrary-length gust sequences and demonstrates broad applicability and robustness across diverse airfoils and inflow conditions.

Accelerating training via expert pretraining and transfer learningDeveloping transformer-based RL for lift control in gusty flowsOvercoming partial observability with limited pressure sensors

Traditional wind farms employ individual turbine control, neglecting wake coupling effects and atmospheric turbulence dynamics, thereby limiting aggregate power output. This paper proposes a reinforcement learning (RL)-based cooperative closed-loop control framework, achieving—for the first time—the real-time integration of RL with high-fidelity large-eddy simulation (LES) to enable dynamic wake-steering decisions. The method incorporates Bayesian optimization to accelerate policy training, overcoming limitations of static, low-fidelity models and open-loop optimization approaches. Experimental results demonstrate a 4.30% increase in total farm power relative to baseline operation—nearly double the 2.19% gain achieved by static yaw optimization. This work establishes the first high-fidelity, dynamic closed-loop paradigm for intelligent cooperative wind farm control.

Enabling closed-loop collaborative wind farm controlIncreasing wind farm power via dynamic wake steeringReplacing static simulators with turbulence-responsive RL

How to craft a deep reinforcement learning policy for wind farm flow control

Jun 06, 2025
EK
Elie Kadoche
🏛️ Polytechnic Institute of Paris | TotalEnergies OneTech | University of Liège

Wind farm power output is significantly reduced by wake effects, and existing control methods lack robustness under time-varying wind conditions. Method: This paper proposes a deep reinforcement learning (DRL)-based flow-field cooperative control strategy tailored for dynamic wind conditions, jointly optimizing yaw angles of individual turbines to mitigate wakes. We introduce a novel DRL architecture integrating Graph Attention Networks (GAT) with multi-head self-attention, coupled with a physics-informed reward function and a progressive training scheme to enhance generalization across arbitrary time-varying wind speeds and directions. Contribution/Results: Experiments demonstrate that the proposed method reduces training steps by ~90% compared to baseline approaches, achieves up to a 14% increase in energy production under dynamic wind conditions, and exhibits substantially improved robustness. This work establishes a scalable, highly adaptive paradigm for real-time intelligent control of large-scale wind farms.

Develop robust wake steering controller for time-varying windsImprove energy production using deep reinforcement learningMitigate wake effects in wind farms for energy optimization

Latest Papers

What's happening recently
View more

This study addresses the challenge of real-time wind field perception and predictive trajectory planning in complex environments, where local airflow is significantly perturbed by surrounding geometry. To this end, we propose WESPR—a lightweight, real-time framework that jointly models environmental geometry and local wind dynamics. By integrating geometric awareness with meteorological data, WESPR rapidly predicts wind-induced disturbances and co-optimizes energy-efficient, safety-aware flight trajectories alongside adaptive control policies. Experimental validation on a Crazyflie quadrotor demonstrates substantial improvements over wind-agnostic adaptive controllers: trajectory deviations are reduced by 12.5%–58.7%, and flight stability increases by 24.6%.

environmental geometryquadrotorreal-time wind prediction

This study addresses the challenge of online energy consumption optimization for data centers within wind farms under uncertainty, where future wind availability and electricity prices are unknown. To enhance the utilization of free wind energy and alleviate the credit assignment problem, the authors propose a reinforcement learning–based scheduling approach that integrates imitation learning with potential-based reward shaping. The method combines Proximal Policy Optimization (PPO) and Soft Actor-Critic (SAC) algorithms equipped with an online update mechanism, enabling wind curtailment–aware workload migration within a reproducible fixed-day simulation framework. Experimental results over a 200-day test set demonstrate that the proposed approach significantly outperforms baseline reinforcement learning strategies. Although it slightly underperforms compared to the offline optimal solution with full foresight, it establishes a robust foundation for extending to multi-site and continuous-time scenarios.

data centersenergy optimizationreinforcement learning

This work addresses the challenge of reliable autonomous navigation for lightweight quadrotors under strong wind disturbances by proposing a hierarchical wind-resilient navigation framework. The upper layer employs deep reinforcement learning (DRL) to generate wind-robust velocity reference trajectories in the inertial frame, while the lower layer utilizes a geometric incremental nonlinear dynamic inversion (INDI) controller for rapid disturbance rejection and high-precision tracking. By incorporating randomized fan-generated wind fields during training, the system generalizes effectively to complex dynamic wind environments without requiring retraining. Flight experiments demonstrate that, under 3.5 m/s wind disturbances, a 50 g quadrotor achieves stable flight at 1.34 m/s, with mission success rate improving from 55.0% to 94.7% and trajectory tracking RMSE reduced by up to 50%.

aerial roboticsautonomous navigationflight stability

This work addresses the challenge of efficiently adapting quadrotor control under nonstationary disturbances such as gusts, payload variations, and ground effects. To this end, we propose a kernel-based domain adaptive control method that leverages random Fourier features to generate diverse disturbances during an offline phase, integrated with differentiable simulation and analytical gradient-based optimization. During online operation, kernel parameters are updated in real time via least-squares estimation, and the kernel bandwidth is adaptively tuned to balance modeling expressiveness with computational efficiency. Requiring only 50 seconds of offline training, the approach enables rapid online adaptation and significantly improves trajectory tracking performance under complex disturbances in both high-fidelity simulation and on the Crazyflie hardware platform, effectively narrowing the sim-to-real gap.

domain-adaptive policy learningnon-stationary disturbancesonline adaptation

Hot Scholars

PA

Pegah Alizadeh

Ericsson Research, Ericsson France
Reinforcement LearningMachine LearningOptimisation for Machine Learning
TM

Thien-Minh Nguyen

Research Asst Prof, NTU Singapore | Lecturer - The University of Queensland (incoming)
Robot Perception and NavigationCooperative RoboticsRobot Learning
LX

Lihua Xie

Professor of Electrical Engineering, Nanyang Technological University
Robust controlNetworked ControlMult-agent Systems
EF

Eduardo F. Camacho

Professor of Automatic Control, University of Seville (Universidad de Sevilla), Spain (España)
ControlModel Predictive Controlsolar energy