Emergent Multi-View Geometry Through Self-Distillation
This study addresses the challenge of disentangling geometry from appearance in multi-view visual representation learning without explicit 3D supervision. We propose Poincaré3, a self-supervised framework that departs from conventional RGB reconstruction paradigms by leveraging masked patch modeling and image-level self-distillation to achieve geometry–appearance disentanglement. Specifically, Poincaré3 introduces a Poincaré adapter that enables a teacher model to observe additional views, facilitating the training of multi-view geometric representations from scratch. Experimental results demonstrate that Poincaré3 significantly outperforms baselines such as DINOv3 across correspondence estimation, camera pose prediction, and 3D reconstruction tasks, while encoding camera motion more precisely. Ultimately, this work establishes a new paradigm for unsupervised multi-view understanding.