Wind-Aware Reinforcement Learning Control of a Small Quadrotor Using Learned Onboard Wind Estimation in Simulated Atmospheric Turbulence

๐Ÿ“… 2026-07-01
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the significant degradation in trajectory tracking performance of small quadrotors operating in the atmospheric boundary layer under strong turbulent winds. To mitigate this, the authors propose a two-stage learning framework: first, an attention-augmented gated recurrent network leverages onboard kinematic and dynamic data to accurately estimate the local three-dimensional wind field, achieving a horizontal wind speed RMSE of 0.40 m/s and a direction error of 3.2ยฐ; second, these wind estimates are integrated into a Proximal Policy Optimization (PPO) reinforcement learning controller to enable wind-aware flight control. This approach represents the first integration of learned wind-field perception with reinforcement learningโ€“based control, reducing trajectory tracking error by 48% compared to a non-wind-aware PD controller across wind speeds of 4โ€“12 m/s, outperforming it in all evaluated scenarios, and maintaining stable flight even under out-of-distribution wind conditions of 13โ€“15 m/s, thereby substantially enhancing wind-robustness.
๐Ÿ“ Abstract
Small multirotor aircraft are increasingly tasked with operations in the atmospheric boundary layer, where turbulent winds comparable to the vehicle's airspeed degrade trajectory tracking and can defeat conventional feedback control. This work illustrates a two-stage learning pipeline that first estimates the local wind from onboard kinematics and dynamics and then exploits that estimate inside a reinforcement learning (RL) flight controller. The wind estimator, an attention-augmented gated recurrent network trained on thousands of simulated flights through von Karman turbulence with power-law shear and veer, recovers the horizontal wind vector with a per-flight root-mean-square error of 0.40 m/s and a direction error of 3.2 degrees on unseen wind regimes, an accuracy near the floor imposed by unresolved turbulence, and generalizes to vertical ascent profiles with a skill score of 0.861 over a constant-wind reference. A proximal policy optimization controller receiving the frozen estimator's output reduces horizontal trajectory tracking error by 48% relative to a wind-blind proportional-derivative baseline across mean winds of 4 m/s to 12 m/s, winning on 100% of evaluation episodes. A three-way ablation decomposes this improvement into a kinematic component, available without wind information, and a wind-perception component; the perception share rises with wind speed, from small in light winds toward roughly half the total benefit in strong winds, consistent with the quadratic scaling of aerodynamic drag. The controller degrades gracefully on out-of-distribution winds of 13 m/s to 15 m/s, where the baseline fails catastrophically.
Problem

Research questions and friction points this paper is trying to address.

quadrotor
atmospheric turbulence
trajectory tracking
wind estimation
flight control
Innovation

Methods, ideas, or system contributions that make the work stand out.

wind estimation
reinforcement learning control
attention-augmented RNN
atmospheric turbulence
quadrotor trajectory tracking
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
A
Abdullah Al Tasim
School of Aerospace and Mechanical Engineering, University of Oklahoma, Norman, Oklahoma, 73019
Wei Sun
Wei Sun
University of Oklahoma
Optimal controldifferential gamesreinforcement learning