🤖 AI Summary
This study addresses the high training costs, low efficiency, and poor generalization of end-to-end methods for vision-based quadrotor navigation through narrow gaps. We propose a two-stage reinforcement learning framework that introduces a novel training architecture combining quasi-analytic policy gradients with critic warm-starting, thereby avoiding full-trajectory backpropagation. The visuomotor policy is further optimized using differentiable simulation, privileged observations, and dual-camera binary mask inputs. Experimental results demonstrate that this approach substantially reduces computational memory overhead while significantly improving sample efficiency and traversal success rates. Real-world flight tests confirm strong robustness to unseen gap geometries and cross-platform generalization capabilities.
📝 Abstract
Traversing narrow gaps is challenging for autonomous quadrotors, especially when control commands come directly from high-dimensional visual observations. Existing end-to-end methods often rely on behavior cloning or full-rollout backpropagation through time (BPTT) via differentiable simulation, which can limit policy performance or incur high training costs. We propose a two-stage reinforcement learning framework for more efficient ego-centric visuomotor gap-traversal policy training, leveraging quasi-analytical policy gradients (QPG) via differentiable simulation and critic warm-starting. The framework utilizes QPG to avoid backpropagation through visual rendering, reducing computation and memory costs while improving sample efficiency. In the first stage, an expert actor and critic are trained using privileged observations, including gap geometry. Unlike prior gap-traversal approaches, our training utilizing QPG does not require resetting the agent along optimized reference trajectories. In the second stage, a visual policy is trained using binary gap masks from two ego-centric cameras and low-dimensional observations, while its privileged critic is warm-started from the first stage. This substantially improves training efficiency and traversal success compared with cold-starting the critic or using full-rollout BPTT. Our framework does not require retraining the expert actor when system parameters change, enabling more efficient generalization across drone platforms than state-of-the-art visual gap-traversal methods based on action supervision. The learned visual policy also generalizes to gaps with unseen shapes. Extensive real-world experiments further demonstrate robust gap traversal using binary masks rendered online. Beyond gap traversal, the proposed framework is generic and can be extended to other visuomotor robot learning tasks.