Learning Agile Gate Traversal via Analytical Optimal Policy Gradient

📅 2025-08-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Agile quadrotor navigation through narrow gates faces challenges including low control precision, complex parameter tuning, poor sample efficiency, and limited interpretability in end-to-end reinforcement learning (RL). Method: This paper proposes a tightly integrated hybrid framework combining neural networks with model predictive control (MPC). It introduces analytical policy gradients to jointly optimize the neural network and an analytical gate-frame detection module, while employing a simplified attitude tracking error representation to ensure training stability. The neural network learns offline reference attitudes and MPC cost weights; MPC then executes high-fidelity trajectory tracking online. Results: Hardware experiments demonstrate rapid, precise gate traversal in constrained environments. The method achieves orders-of-magnitude higher sample efficiency than end-to-end RL, while delivering superior performance, strong interpretability, and robust adaptability to environmental variations.

Technology Category

Multiagent Systems: Multiagent LearningSearch and Optimization: Learning to SearchPlanning, Routing, and Scheduling: Replanning and Plan Repair

Application Category

Responsible Web: Machine-in-the-loop, human agency and autonomySearch and Retrieval-Augmented AI: Agentic searchSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Traversing narrow gates presents a significant challenge and has become a standard benchmark for evaluating agile and precise quadrotor flight. Traditional modularized autonomous flight stacks require extensive design and parameter tuning, while end-to-end reinforcement learning (RL) methods often suffer from low sample efficiency and limited interpretability. In this work, we present a novel hybrid framework that adaptively fine-tunes model predictive control (MPC) parameters online using outputs from a neural network (NN) trained offline. The NN jointly predicts a reference pose and cost-function weights, conditioned on the coordinates of the gate corners and the current drone state. To achieve efficient training, we derive analytical policy gradients not only for the MPC module but also for an optimization-based gate traversal detection module. Furthermore, we introduce a new formulation of the attitude tracking error that admits a simplified representation, facilitating effective learning with bounded gradients. Hardware experiments demonstrate that our method enables fast and accurate quadrotor traversal through narrow gates in confined environments. It achieves several orders of magnitude improvement in sample efficiency compared to naive end-to-end RL approaches.
Problem

Research questions and friction points this paper is trying to address.

Optimizing quadrotor flight through narrow gates
Improving sample efficiency in reinforcement learning
Adaptively tuning MPC parameters via neural network
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hybrid MPC-NN framework for adaptive online tuning
Analytical policy gradients for MPC and gate detection
Simplified attitude error formulation enabling bounded gradient learning
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Tianchen Sun
Tianchen Sun
Zhejiang Lab
Human FactorsIndustrial Engineering
B
Bingheng Wang
Department of Electrical and Computer Engineering, National University of Singapore, Singapore
L
Longbin Tang
Department of Electrical and Computer Engineering, National University of Singapore, Singapore
Y
Yichao Gao
Department of Electrical and Computer Engineering, National University of Singapore, Singapore
L
Lin Zhao
Department of Electrical and Computer Engineering, National University of Singapore, Singapore