Less Decoder is More Encoder: Geometric Representation Learning from Novel View Synthesis

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited geometric representational capacity of encoders in existing novel view synthesis methods, which arises from overly powerful decoders and pixel-level objectives. To this end, we propose SNAP, a self-supervised architecture built upon a Transformer encoder-decoder framework. By constraining decoder expressivity to prevent the suppression of geometric structures, and by introducing pose-conditioned local decoding alongside latent-space reconstruction objectives, SNAP optimizes feature learning and endows patch-level representations with emergent viewpoint invariance. Experimental results demonstrate that SNAP achieves performance comparable to specialized supervised models across five tasks, including localization and pose estimation. Furthermore, it significantly outperforms standard 2D representations under camera displacement scenarios, effectively reducing both computational and data requirements.
📝 Abstract
This paper examines the role of Novel View Synthesis (NVS) in geometric representation learning. In principle, NVS should reason about 3D scene structure, thereby enabling transferable multi-view geometric representations. Yet, existing encoder-based NVS methods yield poor representations. This is not because of a lack of supervisory signal, but rather due to inconspicuous architectural choices: \textit{spatially expressive decoders} that dilute representational capabilities of the scene encoder, and \textit{low-level pixel-space targets} that hinder feature learning. We present SNAP, a self-supervised encoder-decoder transformer that addresses both through a pose-conditioned local decoder and a latent-space reconstruction objective. SNAP is task agnostic, and we show that it is competitive with special-purpose geometry-supervised methods. SNAP also performs competitively against self-supervised representations across five tasks: visual localization, pose estimation, point correspondence, depth estimation, and robot manipulation. Remarkably, SNAP's patch features exhibit emergent viewpoint invariance that approaches heavily supervised models despite lower compute and data budgets. Under camera shifts where standard 2D representations collapse, SNAP degrades more gracefully, revealing that restricting decoder expressivity actively prevents the suppression of transferable geometric structure. https://snap-nvs.github.io
Problem

Research questions and friction points this paper is trying to address.

Novel View Synthesis
Geometric Representation Learning
Self-supervised Learning
Decoder Expressivity
Transferable Representations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Novel View Synthesis
Self-supervised Learning
Geometric Representation Learning
Pose-conditioned Decoder
Latent-space Reconstruction
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Keerthi Kaashyap
Georgia Institute of Technology
D
Dennis Anthony
Georgia Institute of Technology
A
Akshay Krishnan
Georgia Institute of Technology
N
Nhi Ngoc Nguyen
Georgia Institute of Technology
J
Jeremy Collins
Georgia Institute of Technology
James Hays
James Hays
Georgia Tech
Computer VisionRoboticsMachine LearningAI
Shreyas Kousik
Shreyas Kousik
Georgia Institute of Technology
robotics
Animesh Garg
Animesh Garg
Georgia Institute of Technology, University of Toronto
Robotic ManipulationRobot LearningReinforcement LearningMachine LearningComputer Vision