Event3R: Asynchronous-to-Global 3D Reconstruction from Event Camera via Spatial-Temporal Feature Aggregation

πŸ“… 2026-07-17
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the challenges of dense, globally consistent 3D reconstruction from event cameras, which arise from their asynchronous, sparse, and highly dynamic nature, as well as the scarcity of annotated data. To this end, we propose Event3Rβ€”the first end-to-end framework for global 3D reconstruction directly from event streams. Our approach introduces a spatio-temporal voxel representation combined with a temporal attention mechanism to enable time-aware feature fusion. Furthermore, we design a self-supervised pre-training strategy based on Masked Bin Modeling, augmented with contrastive alignment loss and consistency regularization to strengthen cross-view structural correspondence and temporal coherence. Experiments demonstrate that Event3R significantly outperforms existing methods on both synthetic and real-world datasets, achieving robust, temporally consistent, and globally aligned high-quality reconstructions.
πŸ“ Abstract
Robust 3D reconstruction is essential for robotics and embodied perception. Recent feed-forward approaches such as DUSt3R have demonstrated impressive progress in dense 3D reconstruction from RGB images, achieving global geometric consistency and strong generalization. However, extending such dense 3D reconstruction to event cameras remains challenging due to their asynchronous, sparse, and highly dynamic nature, as well as the lack of large-scale, well-labeled datasets. In this work, we introduce Event3R, a feed-forward framework that directly maps asynchronous event streams to globally consistent 3D point clouds. Event3R represents incoming events as spatial-temporal voxels, enabling time-aware feature integration through a temporal attention module that enhances the module's temporal feature learning. To further strengthen temporal representation learning and reduce reliance on labeled data, we propose a Masked Bin Modeling (MBM) strategy for self-supervised pre-training, enabling robust temporal representation learning with minimal labeled data, and retain it as an auxiliary fine-tuning objective. In addition, contrastive alignment and consistency regularization losses are incorporated during fine-tuning to reinforce structural correspondence and temporal coherence across views. Extensive experiments on both synthetic and real-world benchmarks demonstrate that Event3R achieves robust, temporally consistent, and globally aligned 3D reconstructions, significantly outperforming existing event-based methods.
Problem

Research questions and friction points this paper is trying to address.

event camera
3D reconstruction
asynchronous data
temporal consistency
global geometric consistency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Event Camera
3D Reconstruction
Temporal Attention
Masked Bin Modeling
Self-supervised Learning
πŸ”Ž Similar Papers
No similar papers found.