Score
Design algorithms and control policies that estimate scene motion (dense optical flow and visual-inertial odometry), characterize flow uncertainty, and plan active exploratory maneuvers to reveal, localize, and avoid static and dynamic obstacles and to navigate through unknown-shaped gaps using monocular or visual-inertial measurements.
This work addresses the challenge of simultaneously ensuring safe obstacle avoidance and efficient environmental perception for autonomous aerial robots operating in complex, unknown environments. The authors propose an end-to-end reinforcement learning framework that actively controls the onboard camera to jointly optimize navigation and information gathering by fusing robot state, depth images, and local geometric representations. A key innovation lies in embedding a voxelized information metric directly into the reward function, enabling, for the first time, the joint learning of goal-directed motion and exploratory perception, thereby eliciting intrinsic exploratory behavior. Experimental results demonstrate that the proposed method significantly improves both flight safety and environmental exploration efficiency compared to fixed-viewpoint baselines.
This work addresses the degradation of visual-inertial odometry (VIO) performance in visually sparse environments, where autonomous exploration often leads to localization drift and mapping failure. To mitigate this, the authors propose a hierarchical perception-aware exploration framework that explicitly integrates feature quality assessment and continuous yaw trajectory optimization into the exploration strategy. Specifically, a global feature map is constructed to prioritize frontier candidates based on subgoal desirability, while continuous yaw motion is optimized to maintain stable visual feature tracking. Experimental results in both simulated and real-world low-texture environments demonstrate that the proposed method significantly enhances feature tracking stability, reduces odometry drift, and improves exploration coverage efficiency by an average of 30% compared to baseline approaches.
Autonomous obstacle avoidance for monocular optical-flow-driven quadrotors remains challenging due to the lack of differentiable, real-time motion estimation and robust generalization under unseen dense environments. Method: This paper proposes an end-to-end differentiable learning framework comprising (i) a lightweight differentiable optical flow simulator tightly coupled with a simplified quadrotor dynamics model for first-order gradient optimization; (ii) a center-flow attention mechanism to selectively enhance responses to critical ego-motion cues; and (iii) an action-guided active perception strategy to improve cross-scenario robustness. Contribution/Results: Trained exclusively in a minimal simulation environment—without real-world data or domain randomization—the method achieves agile collision-free flight at speeds up to 6 m/s in previously unseen dense scenes. It is successfully deployed on an FPV racing drone platform. Both simulation and physical experiments validate its effectiveness and strong sim-to-real transfer capability.
This work addresses the challenge of autonomous navigation for micro aerial vehicles in complex environments featuring static and dynamic obstacles as well as gaps of unknown geometry. We propose a lightweight, monocular vision–based navigation method that operates without any prior scene knowledge. By modeling optical flow and its associated uncertainty and integrating an active exploration strategy, our approach constructs a real-time navigation stack with minimal computational overhead. Compared to depth-based methods, the proposed system reduces computational cost by several orders of magnitude, enabling deployment on resource-constrained platforms. Real-world experiments demonstrate a 70% success rate across diverse scenarios, achieving performance comparable to depth-based solutions while, to the best of our knowledge, being the first to enable real-time, prior-free traversal of dynamic obstacles and unknown gaps using only a monocular camera.
This work addresses the challenge of severe drift in traditional visual-inertial odometry (VIO) caused by violations of the scene rigidity assumption in non-rigid environments. To this end, we propose DefVINS, a novel framework that explicitly decouples IMU-anchored rigid motion from non-rigid deformations through an embedded deformation graph. A key innovation is the introduction of a visibility-aware conditional activation mechanism that dynamically enables deformation degrees of freedom based on observability analysis. This approach effectively identifies and leverages motion patterns that are otherwise unobservable, significantly enhancing localization robustness in non-rigid scenes while preserving system consistency. Ablation studies validate the efficacy of both the IMU anchoring strategy and the observability-aware deformation activation scheme.
This work addresses the challenges of high-speed FPV quadrotor flight using only a monocular RGB camera in complex environments, where optical flow is often corrupted by mixed motion cues and low signal-to-noise ratios in focus-of-expansion regions. The authors propose decomposing optical flow into translational and rotational components, leveraging only the translational flow—which carries geometric and depth information—and combining it with forward-backward flow inconsistency to generate an uncertainty mask that highlights obstacle structures. This joint representation effectively disentangles ego-motion-induced background flow from obstacle-related flow for the first time, substantially improving perception reliability. An end-to-end neural control policy, trained in a differentiable simulator, achieves robust flight speeds of 13.91 m/s in simulation and 11.79 m/s in real-world forest environments, with a 93.3% success rate over 30 real-world trials—nearly twice the speed of existing comparable systems.
This work addresses the challenge of safe and efficient trajectory planning for unmanned aerial vehicles (UAVs) in unknown, cluttered 3D environments, where limited field-of-view (FOV) and sensing range of onboard sensors hinder reliable navigation. The authors propose a novel approach that directly embeds active perception into trajectory optimization. By leveraging the UAV’s dynamics model, the method accurately encodes FOV geometric constraints in the sensor frame and introduces a velocity-triggered perception mechanism to balance exploration and motion efficiency. It employs parameterized, time-shifted active perception sub-trajectories, enabling online sensing during arbitrary 3D maneuvers without requiring prior maps or dedicated path generators. Built upon a differentiable optimization framework, the planner operates with only a coarse global path as guidance. Extensive simulations and real-world experiments demonstrate the approach’s robustness, safety, and computational efficiency across diverse unknown environments and sensor configurations.
This study addresses the critical need for real-time path adaptation in robotic navigation within dynamic environments, a challenge inadequately covered by existing surveys. Systematically reviewing 138 studies from 2015 to 2025, this work presents the first unified taxonomy of motion planning approaches, categorizing them into sampling-based, graph-search, model predictive control, learning-based, and classical local planners, while integrating both classical and learning-driven methods. It critically examines how dynamic perception influences planning, with in-depth analysis of core challenges including prediction uncertainty, human-robot interaction, and the “freezing robot” problem. The review encompasses key techniques such as velocity obstacles, potential fields, dynamic window approaches, supervised and reinforcement learning, and perception modalities leveraging cameras, LiDAR, and event-based sensors. By establishing a structured methodological framework, this paper offers researchers a comprehensive understanding of the principles, strengths, and limitations across planning paradigms, thereby advancing the field.
本文提出KLTNet,一种基于学习的稀疏特征跟踪器,旨在提高单目视觉惯性里程计在快速运动或低纹理环境下的准确性和鲁棒性。
This study addresses the limitation of existing visual navigation policies, which are constrained by fixed camera configurations and struggle to achieve zero-shot deployment across heterogeneous sensor layouts. To overcome this, we propose a generalizable navigation framework based on explicit geometric projection. Rather than relying on implicit spatial alignment, our method back-projects arbitrary depth sensor data into a unified robot coordinate frame, followed by spherical range-view stitching and validity masking. Combined with aggressive camera randomization during reinforcement learning training, this enables dynamic adaptation to multi-camera setups. Experimental results demonstrate that the proposed approach improves navigation success rates from 78% to 95% under a seven-camera configuration. Furthermore, real-world flight validations confirm its zero-shot transfer capability and robustness against sensor failures.