Think Locally, Refine Globally for Memory-Efficient 3D Reconstruction

📅 2026-09-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
提出LoG-VGGT框架,通过局部跨窗口注意力和全局相机一致性优化,解决长序列3D重建中的内存效率与姿态漂移问题。
📝 Abstract
We propose LoG-VGGT, a memory-efficient framework for long-sequence 3D reconstruction that balances local temporal modeling with global camera consistency. Instead of relying on full global attention, our method introduces cross-window attention at a small subset of transformer blocks, enabling effective information propagation across adjacent temporal windows while keeping memory usage bounded. To mitigate long-term pose drift, we further design a global camera consistency refinement module, where camera tokens interact with compact register tokens via cross-attention to enforce scene-level constraints across the entire sequence. This design enables joint optimization of camera representations and significantly improves long-horizon pose stability without incurring the high cost of sequence-wide attention. Extensive experiments demonstrate that LoG-VGGT achieves improved depth accuracy and robust camera pose estimation across multiple long-sequence benchmarks, while delivering competitive streaming reconstruction performance.
Problem

Research questions and friction points this paper is trying to address.

3D Reconstruction
Memory-Efficient
Long-Sequence
Global Camera Consistency
Innovation

Methods, ideas, or system contributions that make the work stand out.

cross-window attention
global camera consistency refinement
memory-efficient 3D reconstruction
💼 Related Jobs
No related jobs found.
J
Jingke Zhou
Zhejiang University
C
Chenhang Ma
Zhejiang University
Zhizhou Zhong
Zhizhou Zhong
PhD student @ HKUST
face recognitionbiometricsaigc
M
Mingkai Liu
Peking University
Z
Zhuang Zhou
Beijing Institute of Technology
Y
Yicheng Ji
Zhejiang University
B
Binghua Su
Beijing Institute of Technology
B
Bo Cai
Beijing Institute of Technology
X
Xianliang Huang
Fudan University