🤖 AI Summary
Agile quadrotor navigation through narrow gates faces challenges including low control precision, complex parameter tuning, poor sample efficiency, and limited interpretability in end-to-end reinforcement learning (RL).
Method: This paper proposes a tightly integrated hybrid framework combining neural networks with model predictive control (MPC). It introduces analytical policy gradients to jointly optimize the neural network and an analytical gate-frame detection module, while employing a simplified attitude tracking error representation to ensure training stability. The neural network learns offline reference attitudes and MPC cost weights; MPC then executes high-fidelity trajectory tracking online.
Results: Hardware experiments demonstrate rapid, precise gate traversal in constrained environments. The method achieves orders-of-magnitude higher sample efficiency than end-to-end RL, while delivering superior performance, strong interpretability, and robust adaptability to environmental variations.
📝 Abstract
Traversing narrow gates presents a significant challenge and has become a standard benchmark for evaluating agile and precise quadrotor flight. Traditional modularized autonomous flight stacks require extensive design and parameter tuning, while end-to-end reinforcement learning (RL) methods often suffer from low sample efficiency and limited interpretability. In this work, we present a novel hybrid framework that adaptively fine-tunes model predictive control (MPC) parameters online using outputs from a neural network (NN) trained offline. The NN jointly predicts a reference pose and cost-function weights, conditioned on the coordinates of the gate corners and the current drone state. To achieve efficient training, we derive analytical policy gradients not only for the MPC module but also for an optimization-based gate traversal detection module. Furthermore, we introduce a new formulation of the attitude tracking error that admits a simplified representation, facilitating effective learning with bounded gradients. Hardware experiments demonstrate that our method enables fast and accurate quadrotor traversal through narrow gates in confined environments. It achieves several orders of magnitude improvement in sample efficiency compared to naive end-to-end RL approaches.