Less Context, Better Geometry: Masked Geometric Encoder for Robust 3D Foundation Models

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high computational complexity of global attention in 3D foundation models and the geometric instability caused by occlusions and similar views. To this end, we propose a masked geometry encoder coupled with an anchor-guided adaptive token merging mechanism. Specifically, frame tokens are randomly dropped during training while distilling knowledge from a full-context teacher model to enhance single-frame representation robustness. Concurrently, a dynamic token merging algorithm is introduced to substantially reduce computational overhead. This work pioneers a masked geometry encoding strategy that significantly improves reconstruction quality under occlusion. It achieves an excellent trade-off between inference speed and accuracy under limited-view conditions, outperforming existing mainstream methods.
📝 Abstract
Recent progress in 3D foundation models has enabled rapid 3D reconstruction and camera calibration by leveraging learned 3D priors from vast amount of spatial data. However, the all-to-all global attention design leads to quadratic complexity and limits long-sequence inference; unconstrained cross-view interactions also can propagate unreliable evidence from occluded or visually similar but geometrically distant views. In this paper, We introduce a Masked Geometric Encoder (MGE), which promotes the learning of robust geometric representations under incomplete cross-view context. During training, MGE strategically drops frame tokens from global attention and distills from a pretrained full-context teacher model. This allows the model to learn an intrinsically richer per-frame representation while providing sufficient intermediate supervision to avoid performance degradation. Through extensive experiments, we show that MGE leads to much stronger performance under occlusion and doppelganger views while retaining high performance on standard benchmarks. Such a richer frame representation also leads to more effective token reduction during inference. To this end, we develop a novel Anchor-Guided Adaptive token merging technique that preserves representative anchor frames while jointly merging redundant tokens from the remaining views. Compared to other efficient inference approaches, we can achieve inference speedup while consistently maintaining higher reconstruction quality, particularly in limited-view settings.
Problem

Research questions and friction points this paper is trying to address.

3D foundation models
global attention complexity
cross-view interaction
occlusion robustness
token reduction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Masked Geometric Encoder
Knowledge Distillation
Token Merging
3D Foundation Models
Robust 3D Reconstruction
💼 Related Jobs
No related jobs found.
Z
Zhimin Shao
Department of Electrical and Computer Engineering, Johns Hopkins University
X
Xijun Liu
Department of Electrical and Computer Engineering, Johns Hopkins University
Z
Zhaoliang Zhang
Department of Electrical and Computer Engineering, Johns Hopkins University
Y
Yutao Tang
Department of Electrical and Computer Engineering, Johns Hopkins University
A
Abhay Yadav
Department of Electrical and Computer Engineering, Johns Hopkins University
Rama Chellappa
Rama Chellappa
Bloomberg Distinguished Professor, Johns Hopkins University
Image Analysisartificial intelligencebiometricsComputer VisionBiomedical Data Science
C
Cheng Peng
Department of Data Science, University of Virginia