🤖 AI Summary
This study addresses the limited robustness of RGB visual odometry under low-light, motion blur, and frame-drop conditions, as well as the correspondence estimation challenges arising from the asynchronous measurements of event cameras. To overcome these issues, this work proposes an event-based visual odometry system built upon a deep adaptive patch framework. The method introduces dual-modal descriptors that share patch locations, fusing correlation embeddings via learnable scalar gating to independently estimate image and event correspondences while jointly refining ego-motion. Furthermore, a modality-aware keyframe culling strategy preserves sparse constraints, enabling pure-event observation modes. Experimental evaluations on benchmarks such as UZH-FPV demonstrate that the proposed system significantly outperforms baselines like DPVO in trajectory accuracy, maintaining sub-meter precision even under severely degraded conditions.
📝 Abstract
Visual odometry is essential for autonomous navigation in GPS-denied environments, yet RGB-based methods remain vulnerable to motion blur, challenging illumination, and dropped frames. Event cameras complement conventional cameras with high temporal resolution and dynamic range, but their asynchronous measurements complicate reliable correspondence estimation. We present DAPEVO, a learned visual odometry system that estimates image and event correspondences independently at shared patch locations and fuses their correlation evidence before motion refinement. Each tracked patch maintains image and event descriptors, and a learned scalar gate combines modality-specific correlation embeddings for each patch--frame edge before a shared recurrent refinement and bundle-adjustment update. DAPEVO also supports event-only observations, enabling continued tracking when RGB frames are sparse or unavailable, while modality-aware keyframe culling preserves scarce frame constraints. On UZH-FPV, when retaining only one in six RGB frames, DAPEVO's mean absolute trajectory error (ATE) increases by only 36%, from 1.00 to 1.36m, whereas the ATE of DPVO and RAMP-VO rises by factors of $3.7\times$ and $3.1\times$, respectively. On TartanEvent, DAPEVO similarly remains below 1m ATE at 3Hz RGB input, while DPVO and RAMP-VO exceed 9m. Under degraded RGB input on TartanEvent, DAPEVO achieves an ATE of 0.60m, compared with more than 4m for both DPVO and RAMP-VO, while also outperforming event-only DEVO at 0.87m.