🤖 AI Summary
Single-view 3D reconstruction is often limited by the neglect of high-level perceptual information. This work proposes a model-agnostic, plug-and-play perceptual enhancement framework that explicitly integrates semantic and geometric features extracted from pretrained perceptual models into the reconstruction process, enabling perception-guided generation. The method seamlessly integrates into existing 3D reconstruction pipelines and consistently yields significant performance improvements across multiple benchmark datasets when applied to two state-of-the-art approaches, demonstrating the effectiveness and generality of leveraging high-level perceptual signals to enhance reconstruction accuracy.
📝 Abstract
The relationship between object perception and reconstruction is well established in human vision, yet remains underexplored in computer vision. In this paper, we demonstrate that learnt object perception can significantly enhance 3D reconstruction. Focusing on the challenging task of single-view 3D object reconstruction, we propose a method that leverages perceptual signals extracted from pretrained perception models capturing semantic and geometric information to drive the reconstruction of an object from its single image. Our approach is model-agnostic and can be integrated into various reconstruction methods in a plug-and-play manner. Experiments with two state-of-the-art single-view 3D reconstruction pipelines in a benchmark dataset show consistent and substantial improvements achieved by our method, validating the effectiveness of incorporating perception into generation. We provide in-depth analysis of various aspects of our method and its application. Our project page is at https://ynhuhuynh.github.io/perception-3d/.