Battery-Aware Reinforcement Learning for Aggressive Quadrotor Flight

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the thrust degradation caused by battery discharge, which constrains the aggressive flight performance of unmanned aerial vehicles. To overcome this limitation, we propose a voltage-aware deep reinforcement learning framework that, for the first time, jointly trains a load-transient battery model with firmware-level actuator saturation. By employing a voltage feedback feedforward strategy, the approach exploits residual thrust potential while preserving low-level control authority. We demonstrate the critical role of voltage input in enhancing maneuver precision and achieving efficient sim-to-real transfer. Hardware experiments reveal that, compared to a voltage-agnostic baseline, the proposed method reduces trajectory tracking error by 49% and shortens racing lap time by 10.3%, significantly advancing extreme flight performance.
📝 Abstract
Agile flight tasks such as drone racing and pursuit-evasion require strong acceleration and precise turns, but the available thrust changes as the battery discharges and voltage drops under load. Conservative command limits make this variation easier to tolerate, at the cost of unused performance. We investigate how learned controllers can use that additional thrust while retaining the flight controller's voltage compensation and rate control. Our training simulator couples an identified load-transient battery model to rotor dynamics and firmware saturation. The feedforward policy receives filtered voltage during both training and deployment. Controlled ablations distinguish the benefit of a larger thrust-command range from that of voltage information. On a 38 g Crazyflie Brushless, the resulting policy reduces circle tracking error by 49% relative to stock-authority RL at 3.84 m/s, while preserving easy-task precision. Mean 20-lap race time decreases from 106.22 s to 95.24 s. Compared with a voltage-blind policy with the same increased authority, hardware error and race time are lower by 15.3% and 4.5%, respectively. In simulation, replacing the policy's voltage input with a recording from a different battery condition worsens hard-circle tracking, with a smaller, voltage-dependent effect in racing. Together, these results show where a simple voltage input complements existing actuator compensation in aggressive learned flight.
Problem

Research questions and friction points this paper is trying to address.

agile quadrotor flight
battery voltage drop
thrust authority
reinforcement learning
drone racing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Battery-aware reinforcement learning
Aggressive quadrotor flight
Load-transient battery model
Voltage compensation
Feedforward policy
💼 Related Jobs
No related jobs found.
A
Alejandro Sanchez Roncero
Department of Robotics, Perception and Learning, School of Electrical Engineering and Computer Science, Royal Institute of Technology (KTH), SE-100 44 Stockholm, Sweden
Olov Andersson
Olov Andersson
Assistant Professor at KTH Royal Institute of Technology. Previously: ASL@ETH Zurich
Robot LearningAutonomous RobotsMotion PlanningMappingNavigation
P
Petter Ogren
Department of Robotics, Perception and Learning, School of Electrical Engineering and Computer Science, Royal Institute of Technology (KTH), SE-100 44 Stockholm, Sweden