Score
Design and implement mathematical representations and algorithms that encode projective relationships between cameras and image planes, using homogeneous coordinates and algebraic pairwise projective relations. Build components that compute and apply epipolar entities (fundamental/essential matrices and epipolar lines), constrain correspondence search to epipolar manifolds, transfer heatmaps or features between views, and produce linearized, attention‑compatible geometric embeddings and camera models (e.g., RPC/RFM) for integration into matching, reconstruction, or attention‑based pipelines.
This paper addresses the fragmented and unsystematic modeling of geometric constraints in deep learning by proposing the first unified taxonomy of geometric constraints tailored for modern deep learning frameworks. Methodologically, it systematically integrates multi-view geometry, epipolar constraints, camera calibration models, self-supervised geometric consistency losses, and differentiable rendering to establish a three-dimensional classification framework spanning modeling principles, integration strategies, and optimization objectives. The contributions are threefold: (1) clarifying the applicability boundaries and failure mechanisms of over one hundred geometric constraints across vision tasks such as depth estimation; (2) uncovering key design paradigms for synergistic co-design of geometric priors and neural architectures; and (3) identifying principled pathways to overcome three core challenges—dynamic scenes, textureless regions, and cross-domain generalization.
This work proposes a pixel-accurate epipolar-guided matching method to address the limitations of conventional approaches in challenging scenarios such as repetitive textures or large baselines, where coarse spatial binning introduces errors, necessitates post-processing, and often misses valid correspondences. By leveraging the fundamental matrix to enforce epipolar geometry, the method defines for each keypoint an angular interval derived from a tolerance circle in angle space, transforming the matching problem into a one-dimensional interval query. Efficient exact matching is achieved via a segment tree with logarithmic time complexity. The approach enables per-point tolerance control, eliminates approximation errors and redundant descriptor comparisons, and recovers a complete set of correspondences without post-processing. Evaluated on the ETH3D dataset, it significantly outperforms existing methods, achieving both higher completeness and notable speedup.
This paper addresses the solvability problem of viewing graphs in structure-from-motion (SfM), i.e., under what conditions epipolar constraints among camera pairs uniquely determine all camera poses. Tackling a long-standing conjecture, we introduce algebraic geometry—integrating projective geometry, graph theory, and polynomial ideal theory—to establish a rigorous solvability framework grounded in the dimension of algebraic varieties. Our approach yields necessary and sufficient algebraic conditions for solvability and provides a decidable verification procedure for arbitrary viewing graphs. This constitutes the first theoretically complete analytical tool for topology-aware SfM system design, enabling principled assessment of camera network connectivity and geometric uniqueness.
Limited 3D perception in multi-view vision tasks stems from insufficient camera geometry modeling. To address this, we propose Projective Positional Encoding (PRoPE), the first method to encode the full camera intrinsic and extrinsic parameters—defining the frustum geometry—as relative positional encodings within Transformers. PRoPE jointly integrates token-level ray-map encoding, attention-level relative pose encoding, and geometrically grounded positional encoding to explicitly model cross-view spatial relationships in self-attention. Crucially, it supports generalization across varying sequence lengths, diverse intrinsic parameter distributions, and out-of-distribution (OOD) camera configurations. Extensive experiments on multi-view image synthesis and stereo depth estimation demonstrate consistent performance gains across model scales; improvements are especially pronounced for long sequences, unseen intrinsics, and OOD scenarios. These results validate the broad efficacy of geometry-aware positional encoding for enhancing multi-view Transformers.
This paper addresses the relative pose estimation problem for three calibrated cameras given only four correspondences across all views. To overcome limitations of conventional methods—namely, their reliance on more correspondences or insufficient robustness—we propose a novel strategy that approximates a fifth correspondence using the centroid of the four observed points. We further introduce the first joint three-view pose estimation framework integrating a 4-point affine fundamental matrix solver, a standard 5-point relative pose solver, and a P3P solver. Geometric modeling enhances robustness against noise and outliers, while local optimization refines accuracy. Evaluated on real-world datasets, our method achieves state-of-the-art performance: the centroid-based strategy significantly outperforms pure affine approaches, striking a superior balance among accuracy, robustness, and computational efficiency, with straightforward implementation.
This work addresses the limitations of existing feedforward view synthesis methods, which rely on Plücker ray representations that are highly sensitive to camera coordinate systems, resulting in poor cross-view geometric consistency. To overcome this, the authors propose a projection-conditioning strategy that replaces raw ray inputs with 2D projection cues from the target view, effectively reformulating the task as a stable image-to-image translation problem. A tailored masked autoencoder pretraining mechanism is introduced to leverage large-scale uncalibrated data under this new conditioning paradigm. The proposed approach significantly enhances model robustness and view consistency, achieving state-of-the-art performance across multiple novel view synthesis benchmarks. Notably, it outperforms ray-based baselines by a clear margin on geometric consistency metrics, demonstrating the effectiveness of decoupling geometry representation from explicit ray parameterization.
This work proposes a novel method for establishing point correspondences across image sequences in real time under unknown 3D scene structure and imaging geometry. The approach introduces a channel-vector-based uncertainty density model and employs an online optimization mechanism driven by Neyman chi-square divergence to iteratively learn mappings between image point sets. By representing channel vectors with basis functions and integrating a density divergence criterion, the algorithm achieves rapid convergence and high-accuracy correspondence estimation under general imaging geometries. Experimental results demonstrate that the proposed method outperforms state-of-the-art techniques across multiple metrics, offering a compelling combination of real-time performance, robustness, and accuracy.
This work addresses the physical inconsistency in conventional multi-view satellite image evaluation, which relies on unconstrained 2D matching and ignores the epipolar geometry implicitly encoded in Rational Polynomial Coefficients (RPCs). The paper proposes the first geometry-aware evaluation protocol tailored to the RPC framework: it constructs a geometrically constrained search manifold via 3D projection and employs dense matching as a proxy task to assess the local uniqueness of features within a physically plausible space. By integrating geometric constraints into foundational model evaluation for the first time, this approach reveals a decoupling between semantic consistency and geometric localization capability, and establishes a reproducible, geometry-faithful benchmark for satellite imagery. Experiments demonstrate that, under RPC-consistent evaluation, generic 2D backbone networks outperform specialized 3D-aware models, underscoring the fundamental importance of geometric constraints in task formulation.
本文提出一种统一框架,通过2D重投影测量优化几何特征和高阶关系,解决了在保持计算效率的同时丰富3D结构的问题。
Existing linear probes struggle to uncover the internal encoding structure of geometric information in self-supervised vision Transformers (ViTs). This work proposes a controlled subspace intervention framework that leverages singular value decomposition (SVD) on converged linear probe weights to isolate a low-rank subspace carrying explicit geometric signals. For the first time, subspace analysis reveals distinct differences in geometric representation between DINOv2 and MAE, demonstrating that geometric information is highly compressible, peaks in accuracy at intermediate network layers, and exhibits pronounced low-rank characteristics. These findings provide both theoretical grounding and practical design guidance for lightweight decoders and efficient feature selection strategies in self-supervised vision models.
We present an efficient and robust method for 3D geometric reconstruction that is based solely on the camera-independent linear relationships among a given set of points, which are stable over time and robustly estimated using multiple point matches. We essentially learn, from correspondences between points across several frames, a linear geometric auto-regression matrix $\mathbf{W}$, which establishes how a point in 3D can be expressed as a linear combination of all the others. This matrix is constant and does not depend on the world coordinate system or the camera pose---it is an intrinsic property of the point set. We also show that the principal eigenvectors of $\mathbf{W}$, which all have eigenvalue $1$, provide a homogeneous representation of the 3D point configuration. The first version of our method takes advantage of noisy monocular depth maps in order to obtain, from multiple frames, a robust geometric auto-regression matrix $\mathbf{W}$ of linear relationships between the 3D points. Thus, we build on recent advances in deep learning, which now provide monocular depth estimation models that are fast but very often noisy. Our approach handles noise through robust linear estimation over several frames. The second version of our method does not need monocular depth estimation maps. It applies in cases of weak-perspective projection, when the linear combinations between the 3D points can be robustly estimated from their 2D projections in the image. Note that the camera projection matrix is never used in our derivations. Consequently, our method does not recover camera pose, but only 3D structure. This is a key difference between our method and the related literature on 3D geometric reconstruction.