Score
Design and implement modules and estimation methods that predict geometric transformation parameters (e.g., affine or rigid) and apply them to visual data representations; examples include networks that estimate transforms and warp image regions or feature maps to canonical poses or stabilized coordinates. Build algorithms to estimate frame‑to‑frame motion, compensate for motion blur and occlusion, and stabilize or align features over time to improve spatial grounding and tracking robustness.
This paper addresses the fragmented and unsystematic modeling of geometric constraints in deep learning by proposing the first unified taxonomy of geometric constraints tailored for modern deep learning frameworks. Methodologically, it systematically integrates multi-view geometry, epipolar constraints, camera calibration models, self-supervised geometric consistency losses, and differentiable rendering to establish a three-dimensional classification framework spanning modeling principles, integration strategies, and optimization objectives. The contributions are threefold: (1) clarifying the applicability boundaries and failure mechanisms of over one hundred geometric constraints across vision tasks such as depth estimation; (2) uncovering key design paradigms for synergistic co-design of geometric priors and neural architectures; and (3) identifying principled pathways to overcome three core challenges—dynamic scenes, textureless regions, and cross-domain generalization.
This work addresses the coverage failure of conformal prediction (CP) under geometric distribution shifts—such as rotations and reflections—where standard CP guarantees degrade. We propose a pose-normalization-augmented CP framework, whose core innovation is the first integration of pose normalization as a geometry-aware feature extractor within the CP pipeline. Crucially, it requires no modification to the underlying black-box predictor and uniformly handles both discrete and continuous geometric transformations while preserving rigorous marginal coverage guarantees. By modeling geometric invariance through normalized features and adapting CP via a standardized interface, our method enhances robustness without compromising CP’s formal statistical assurances. Experiments demonstrate stable empirical coverage ≥95% across diverse geometric shifts, significantly outperforming equivariant models and data-augmentation baselines. Moreover, the framework is fully compatible with arbitrary pre-trained predictors.
This study addresses the problem of improving accuracy in 3D reconstruction and surface normal estimation in stereo vision by leveraging affine correspondences. Recognizing that conventional methods often neglect local geometric deformations, the authors propose a novel approach to estimate local affine transformations from oriented image correspondences and integrate them into the fundamental matrix estimation and epipolar geometry framework to enhance surface normal reconstruction. To quantitatively evaluate performance, a specialized calibration object comprising three mutually orthogonal checkerboard planes is constructed, and experiments are conducted on both synthetic and real images. Results demonstrate that, under typical stereo configurations and planar orientations, the proposed method achieves surface normal estimation errors of only a few degrees in real-world scenes, thereby validating the efficacy and practical limits of modeling affine correspondences.
This work addresses the insufficient integration of learning-based methods and geometric constraints in camera pose and scene structure estimation by proposing a modular framework. The approach first employs a learning model (VGGT) to generate initial hypotheses for depth and relative pose, which are subsequently refined and validated using classical geometric algorithms such as point-to-plane RGB-D ICP. Crucially, the framework explicitly distinguishes the roles of learning as a “proposer” and geometry as a “referee,” emphasizing that the geometric module serves not merely as post-processing but as an essential mechanism for verifying and integrating learned outputs. Experiments on the TUM RGB-D dataset demonstrate that, in moderately challenging rigid scenes, the system significantly outperforms both purely learning-based and purely geometric baselines when the learned depth aligns geometrically with the camera intrinsics and undergoes optimization by the geometric backend.
While existing multi-frame models achieve cross-frame consistency, their single-frame accuracy often lags behind that of single-frame methods. Through systematic ablation studies, this work demonstrates that data diversity and quality are critical for 3D geometry estimation and reveals that commonly used loss functions may inadvertently suppress performance. To address these issues, the authors propose CARVE, a novel approach integrating a high-resolution network architecture, joint sequence- and frame-level supervision, a consistency loss, and alignment between depth maps and camera parameters. CARVE achieves state-of-the-art and robust performance across multiple benchmarks in tasks including point cloud reconstruction, video depth estimation, and estimation of camera pose and intrinsics.
In large-scale Structure-from-Motion (SfM), sparse inter-view overlap and drastic viewpoint changes—especially in aerial-to-ground scenarios—lead to low cross-image feature matching density and weak geometric consistency. To address this, we propose a geometry-guided hybrid matching paradigm: (1) geometric verification is formulated as an optimization problem based on Sampson distance; (2) detector-agnostic dense matching is fused with detector-driven sparse anchor guidance, where sparse anchors constrain and enhance the geometric consistency of dense matches; and (3) multi-view geometric consistency is explicitly modeled. Our method significantly improves both matching density and accuracy, outperforming state-of-the-art approaches in extreme large-scale settings. Consequently, camera pose estimation becomes more accurate, and the reconstructed 3D point cloud achieves higher completeness and fidelity.
This work addresses the common trade-off in efficient camera pose estimation, where speed is often achieved at the expense of accuracy, and proposes a hybrid paradigm that integrates neural network-based initializations with classical Structure-from-Motion (SfM) optimization. The approach maintains high reconstruction accuracy while significantly reducing the number of required feature points. To systematically evaluate the efficacy of sparse matching and neural initialization in guiding bundle adjustment, the authors construct a novel SfM benchmark tailored for novel view synthesis. Experiments demonstrate that merely lowering feature density can accelerate conventional SfM pipelines, yet the combination of neural priors with traditional optimization achieves the best balance between efficiency and accuracy. The publicly released benchmark aims to advance research in high-precision, efficient SfM methods.
This work addresses the limitations of traditional 3D reconstruction methods, which predict point maps in camera-centered coordinates, struggle to incorporate scene structural priors, and suffer from high rotational degrees of freedom across views, leading to inconsistent reconstructions. To overcome these issues, the authors propose predicting point maps in a gravity-aligned upright coordinate system, thereby reducing inter-view rotational ambiguity through a shared vertical axis. They introduce the Gravity Grounded Geometry Transformer (G3T) model and the G3T-Long incremental reconstruction framework, which for the first time integrate gravity-aligned coordinates into point map prediction by combining a Transformer architecture, gravity-aware pose estimation, and a submap stitching strategy. Experiments demonstrate that this approach significantly improves reconstruction accuracy and robustness, outperforming existing methods in incremental 3D reconstruction and validating the effectiveness of gravity-aligned representations.
This study addresses the limited geometric understanding of current vision–language–action (VLA) models, which constrains their performance in embodied tasks. For the first time, the authors quantify the “geometry gap” between VLAs and geometric foundation models (GFMs) using linear probing, and systematically evaluate—under a unified experimental setup—the impact of three fusion architectures, training data scale, and multi-view inputs on geometric perception. The findings reveal that specific fusion architectures substantially enhance geometric comprehension, while multi-view observations and sufficient training data are critical for high performance. This work establishes key design principles and provides empirical evidence for developing geometry-aware VLA systems.
This work addresses the challenge of simultaneously achieving high accuracy, robustness, and loop-closure capability in two-frame pose optimization by proposing a unified framework that integrates geometric and photometric information. For the first time, dense geometric feature descriptors are incorporated into differential photometric optimization, replacing conventional photometric residuals with descriptor-based residuals to enable subpixel-level pose estimation in descriptor space. By synergistically combining the strengths of both geometric and photometric paradigms, this approach explores a novel trajectory for pose optimization grounded in descriptor similarity. Experimental results demonstrate a significant improvement in tracking accuracy; however, overall performance remains slightly inferior to reprojection error–based methods, with the primary bottleneck identified as the relatively flat landscape of descriptor similarity, which limits optimization efficacy.