Learning Vision-Based Agile Gap Traversal: Differentiable Simulation with a Warm-Started Critic

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high training costs, low efficiency, and poor generalization of end-to-end methods for vision-based quadrotor navigation through narrow gaps. We propose a two-stage reinforcement learning framework that introduces a novel training architecture combining quasi-analytic policy gradients with critic warm-starting, thereby avoiding full-trajectory backpropagation. The visuomotor policy is further optimized using differentiable simulation, privileged observations, and dual-camera binary mask inputs. Experimental results demonstrate that this approach substantially reduces computational memory overhead while significantly improving sample efficiency and traversal success rates. Real-world flight tests confirm strong robustness to unseen gap geometries and cross-platform generalization capabilities.
📝 Abstract
Traversing narrow gaps is challenging for autonomous quadrotors, especially when control commands come directly from high-dimensional visual observations. Existing end-to-end methods often rely on behavior cloning or full-rollout backpropagation through time (BPTT) via differentiable simulation, which can limit policy performance or incur high training costs. We propose a two-stage reinforcement learning framework for more efficient ego-centric visuomotor gap-traversal policy training, leveraging quasi-analytical policy gradients (QPG) via differentiable simulation and critic warm-starting. The framework utilizes QPG to avoid backpropagation through visual rendering, reducing computation and memory costs while improving sample efficiency. In the first stage, an expert actor and critic are trained using privileged observations, including gap geometry. Unlike prior gap-traversal approaches, our training utilizing QPG does not require resetting the agent along optimized reference trajectories. In the second stage, a visual policy is trained using binary gap masks from two ego-centric cameras and low-dimensional observations, while its privileged critic is warm-started from the first stage. This substantially improves training efficiency and traversal success compared with cold-starting the critic or using full-rollout BPTT. Our framework does not require retraining the expert actor when system parameters change, enabling more efficient generalization across drone platforms than state-of-the-art visual gap-traversal methods based on action supervision. The learned visual policy also generalizes to gaps with unseen shapes. Extensive real-world experiments further demonstrate robust gap traversal using binary masks rendered online. Beyond gap traversal, the proposed framework is generic and can be extended to other visuomotor robot learning tasks.
Problem

Research questions and friction points this paper is trying to address.

vision-based gap traversal
autonomous quadrotors
end-to-end visuomotor control
training efficiency
policy generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Differentiable Simulation
Quasi-Analytical Policy Gradients
Critic Warm-Starting
Visuomotor Reinforcement Learning
Agile Gap Traversal
N
Nuthasith Gerdpratoom
Department of Electrical and Computer Engineering, National University of Singapore, 4 Engineering Drive 3, Singapore 117583, Singapore
Tianchen Sun
Tianchen Sun
Zhejiang Lab
Human FactorsIndustrial Engineering
Y
Yichao Gao
Department of Electrical and Computer Engineering, National University of Singapore, 4 Engineering Drive 3, Singapore 117583, Singapore
Lin Zhao
Lin Zhao
Assistant Professor, National University of Singapore
control theoryreinforcement learningroboticsautonomous vehiclespower system