🤖 AI Summary
This study addresses the limitation of existing visual navigation policies, which are constrained by fixed camera configurations and struggle to achieve zero-shot deployment across heterogeneous sensor layouts. To overcome this, we propose a generalizable navigation framework based on explicit geometric projection. Rather than relying on implicit spatial alignment, our method back-projects arbitrary depth sensor data into a unified robot coordinate frame, followed by spherical range-view stitching and validity masking. Combined with aggressive camera randomization during reinforcement learning training, this enables dynamic adaptation to multi-camera setups. Experimental results demonstrate that the proposed approach improves navigation success rates from 78% to 95% under a seven-camera configuration. Furthermore, real-world flight validations confirm its zero-shot transfer capability and robustness against sensor failures.
📝 Abstract
Existing visual navigation policies are inherently bound to fixed camera configurations, creating a fundamental barrier to zero-shot deployment across heterogeneous robot sensor layouts. To overcome this limitation, we present an embodiment-informed navigation policy capable of generalizing across diverse depth sensor configurations on a specific aerial platform. Instead of implicitly learning spatial alignments, our approach explicitly unprojects depth measurements from arbitrary depth sensor payloads, varying in sensor count, mounting extrinsics, and intrinsics, into a shared robot-centric frame, stitching them into a unified spherical range image and a binary validity mask. This mask allows the downstream policy to explicitly distinguish covered space from unobserved blind spots. Trained via reinforcement learning with aggressive camera randomization, our policy generalizes zero-shot to unseen layouts featuring up to seven cameras, scaling success rates from 78% to 95% as total spatial sensing coverage increases. Finally, real-world flight trials on a physical quadrotor, conducted in an obstacle-filled corridor and an outdoor forest, validate the policy's zero-shot transfer across camera configurations and its resilience to sudden online sensor dropouts.