🤖 AI Summary
To address the challenge of image feature matching without pixel-level annotations or prior knowledge of camera poses/depth, this paper proposes an instructional learning (IL) self-supervised paradigm. Methodologically, correspondence learning is formulated as a bilevel optimization problem, enabling—for the first time—the differentiable implementation of bundle adjustment (BA): the external BA reprojection error serves as an implicit supervisory signal, and gradients are backpropagated via implicit differentiation to jointly optimize feature matching and geometric consistency in an end-to-end manner. The framework requires neither explicit annotations nor pose labels, relying solely on geometric constraints inherent in video sequences. On both feature matching and relative pose estimation tasks, it achieves a 30% average accuracy improvement over state-of-the-art unsupervised methods, establishing a novel paradigm for unsupervised correspondence learning.
📝 Abstract
Learning feature correspondence is a foundational task in computer vision, holding immense importance for downstream applications such as visual odometry and 3D reconstruction. Despite recent progress in data-driven models, feature correspondence learning is still limited by the lack of accurate per-pixel correspondence labels. To overcome this difficulty, we introduce a new self-supervised scheme, imperative learning (IL), for training feature correspondence. It enables correspondence learning on arbitrary uninterrupted videos without any camera pose or depth labels, heralding a new era for self-supervised correspondence learning. Specifically, we formulated the problem of correspondence learning as a bilevel optimization, which takes the reprojection error from bundle adjustment as a supervisory signal for the model. To avoid large memory and computation overhead, we leverage the stationary point to effectively back-propagate the implicit gradients through bundle adjustment. Through extensive experiments, we demonstrate superior performance on tasks including feature matching and pose estimation, in which we obtained an average of 30% accuracy gain over the state-of-the-art matching models.