🤖 AI Summary
This study addresses the challenge of disentangling geometry from appearance in multi-view visual representation learning without explicit 3D supervision. We propose Poincaré3, a self-supervised framework that departs from conventional RGB reconstruction paradigms by leveraging masked patch modeling and image-level self-distillation to achieve geometry–appearance disentanglement. Specifically, Poincaré3 introduces a Poincaré adapter that enables a teacher model to observe additional views, facilitating the training of multi-view geometric representations from scratch. Experimental results demonstrate that Poincaré3 significantly outperforms baselines such as DINOv3 across correspondence estimation, camera pose prediction, and 3D reconstruction tasks, while encoding camera motion more precisely. Ultimately, this work establishes a new paradigm for unsupervised multi-view understanding.
📝 Abstract
Over a century ago, Henri Poincaré argued that a motionless observer cannot acquire the notion of space. Yet, most visual representation learning methods operate on individual images, while those that leverage multiple views rely on RGB reconstruction, entangling geometry with appearance. We propose Poincar3, a self-supervised method that learns representations from multiple views through self-distillation instead of RGB reconstruction. We combine masked patch and image-level distillation with a teacher that observes additional views, enabling training from scratch without explicit 3D supervision. Poincar3 outperforms both previous single and multi-view self-supervised approaches such as DINOv3, MuM, and Muskie on correspondence estimation, camera pose estimation, and 3D reconstruction. Using a lightweight Poincaré adapter, we also find that our learned features encode camera motion more accurately than existing self-supervised representations.