Score
Designs and evaluates multi-stage models or pipelines that decompose the mapping between image pixels and world-coordinate representations into explicit substeps, including both forward and reverse (bidirectional) mappings. Builds per-step supervised mappings and consistency constraints that enforce joint agreement across coordinate frames and substeps to improve geometric and multi-task prediction quality.
Traditional SfM/MVS methods suffer from complex pipelines, high computational cost, and poor robustness in textureless regions. To address these limitations, this paper systematically reviews feedforward, single-pass 3D reconstruction techniques—particularly end-to-end models such as DUSt3R—and proposes a Transformer-based paradigm for joint pose and geometry estimation, eliminating iterative optimization and unifying multi-view geometry and camera pose estimation. The method supports arbitrary numbers of input images, ensuring strong generalization and flexibility. By integrating correspondence modeling, multi-view feature fusion, and joint regression, it achieves superior efficiency and robustness over both classical approaches and learning-based methods (e.g., MVSNet) on standard benchmarks. We further survey prevailing architectural frameworks, benchmark datasets, and evaluation metrics, and identify key future directions—including dynamic scene modeling and model scalability.
This paper addresses the fragmented and unsystematic modeling of geometric constraints in deep learning by proposing the first unified taxonomy of geometric constraints tailored for modern deep learning frameworks. Methodologically, it systematically integrates multi-view geometry, epipolar constraints, camera calibration models, self-supervised geometric consistency losses, and differentiable rendering to establish a three-dimensional classification framework spanning modeling principles, integration strategies, and optimization objectives. The contributions are threefold: (1) clarifying the applicability boundaries and failure mechanisms of over one hundred geometric constraints across vision tasks such as depth estimation; (2) uncovering key design paradigms for synergistic co-design of geometric priors and neural architectures; and (3) identifying principled pathways to overcome three core challenges—dynamic scenes, textureless regions, and cross-domain generalization.
Existing 3D reconstruction methods struggle to balance generalization and practicality due to either inefficient per-scene optimization or reliance on category-specific training. This work proposes a feed-forward, output-representation-agnostic framework for 3D reconstruction, systematically addressing five core challenges: feature enhancement, geometry awareness, model efficiency, data augmentation, and temporal modeling. By unifying the analysis of image backbones, multi-view fusion mechanisms, and geometric priors—and integrating major datasets and evaluation benchmarks—it establishes a standardized benchmarking protocol. The study transcends differences in geometric representations, formulates a problem-driven, generalizable modeling paradigm, and outlines promising future directions in scalability, evaluation metrics, and world modeling.
Existing CAD learning approaches discretize B-Rep models into triangle meshes, thereby discarding the analytical surface representations and topological information essential for consistent instance-level analysis. This work proposes STEP-Parts, a deterministic pipeline that directly extracts geometric instance partitions from native STEP B-Rep data. The method defines partitions based on intrinsic B-Rep topology, merges faces using analytical surface types and near-tangent plane continuity criteria, and transfers labels to triangulated meshes via face-to-mesh correspondence mapping. STEP-Parts ensures boundary consistency across varying triangulations and processes the DeepCAD subset of the ABC dataset—comprising approximately 180,000 models—in under six hours. The resulting labels significantly enhance performance in implicit reconstruction-segmentation tasks and point cloud networks. Code and precomputed labels are publicly released.
The field of 3D vision suffers from fragmented data representations, learning paradigms, and benchmarking protocols, leading to a lack of unified understanding regarding efficiency, fidelity, and scalability. This work proposes the first cohesive conceptual framework that integrates geometric representations—such as point clouds, meshes, voxels, and 3D Gaussians—with diverse learning paradigms—including 2D-supervised learning, implicit neural representations, and 4D modeling—and connects them to real-world application scenarios. By constructing a structured knowledge graph of 3D vision, the study systematically relates dataset design, supervision mechanisms, and task requirements, clarifying the trade-offs between efficiency and fidelity and charting pathways for multimodal geometric grounding. This framework offers systematic guidance for reconstruction, generation, and dynamic scene modeling, advancing the field toward a unified and efficient paradigm.
This work addresses the problem of efficiently merging multiple fine-tuned models into a unified multitask model without retraining. The authors formalize model merging as a convex quadratic program over residual updates, achieving theoretically optimal fusion by calibrating inputs and outputs to minimize calibration error in the output space. This study provides the first formal optimality guarantees for model merging, introduces an interpretable diagnostic metric based on residual energy, and unifies existing heuristic approaches within a single theoretical framework as special cases. Experimental results demonstrate that the proposed method matches or surpasses current techniques in single-layer settings and consistently improves performance across language and vision benchmarks in multilayer merging scenarios. Furthermore, the quality of merged models can be accurately predicted using a small calibration set.
Existing image warping methods require task-specific model training, exhibiting poor generalization and limited adaptability to diverse camera models or custom distortions. To address this, we propose MOWA—a unified, multi-task image warping model capable of handling six distinct warping tasks within a single architecture. Our approach introduces three key innovations: (1) a region-pixel two-level motion disentanglement mechanism for enhanced geometric modeling accuracy; (2) a lightweight point-based classifier that dynamically generates task-aware prompts for conditional feature modulation; and (3) end-to-end differentiable warping networks jointly optimized with multi-scale motion estimation. Extensive experiments demonstrate that MOWA consistently outperforms dedicated state-of-the-art models across all six tasks. Moreover, it exhibits strong cross-camera generalization and zero-shot transfer capability, enabling robust adaptation without task-specific retraining.
While existing multi-frame models achieve cross-frame consistency, their single-frame accuracy often lags behind that of single-frame methods. Through systematic ablation studies, this work demonstrates that data diversity and quality are critical for 3D geometry estimation and reveals that commonly used loss functions may inadvertently suppress performance. To address these issues, the authors propose CARVE, a novel approach integrating a high-resolution network architecture, joint sequence- and frame-level supervision, a consistency loss, and alignment between depth maps and camera parameters. CARVE achieves state-of-the-art and robust performance across multiple benchmarks in tasks including point cloud reconstruction, video depth estimation, and estimation of camera pose and intrinsics.
This work addresses the high computational cost of high-fidelity 3D generation, which typically relies on large-scale data and models while underutilizing the rich semantic and structural priors embedded in discriminative 3D foundation models. To bridge this gap, the authors propose ROAD, a novel framework that, for the first time, transfers priors from discriminative 3D foundation models into a diffusion Transformer. ROAD introduces a reciprocal objective alignment mechanism that effectively reconciles the heterogeneity between generative and discriminative latent spaces through global semantic compression and optimal micro-structure matching—formulated as bipartite graph matching. Remarkably, without increasing inference overhead, ROAD achieves generation quality on par with the industrial baseline Step1X-3D using only 1.5% of the training data, substantially reducing both training cost and computational requirements.
This work addresses the challenges of missing correspondences and incomplete geometric information in point cloud reconstruction from partially observed multi-view inputs. The authors propose a training-free optimization method that jointly recovers the 3D point cloud and its cross-view projection mappings. Built upon an extended multi-view synchronized embedding framework, the approach integrates variable projection, geometric constraints, and visibility modeling, making it applicable to both fixed and variable projection settings without requiring category-specific priors. Experiments on ShapeNet and Pix3D demonstrate that the method robustly reconstructs partial multi-view point clouds, consistently outperforming existing non-learning baselines across Chamfer distance, Earth Mover’s Distance (EMD), and Reconstruction Overlap Accuracy (ROA) metrics.
Traditional visual navigation struggles to balance global geometric consistency with topological generalization, limiting its performance in complex environments. This work proposes a novel map representation based on pixel-level relative 3D connectivity, which constructs a pixel correspondence graph in a relative coordinate frame from image sequences and generates a “WayPixel Costmap” for planning and control. By preserving high-fidelity geometric information without requiring global geometric consistency, the approach overcomes the limitations of conventional topological graphs and dense reconstructions. Experimental results demonstrate that the method significantly outperforms image-level and object-level representations across four simulated tasks and real-world scenarios, validating its accuracy and practicality for visual navigation.
This work addresses the limitations of existing panoramic stitching methods, which rely on pairwise feature matching and often fail to maintain multi-view geometric consistency in complex scenes characterized by weak textures, large disparities, or repetitive patterns, leading to misalignments and distortions. To overcome these challenges, the authors propose a photogrammetry-driven global alignment framework that leverages estimated camera poses to align images in 3D space. They introduce a novel 3D-aware Transformer architecture that explicitly models multi-view geometric consistency through joint feature optimization and cross-view information aggregation. Key contributions include the first formulation of multi-view consistency in 3D space, a Transformer-based 3D-aware stitching network, and the creation of the first large-scale real-world panoramic stitching dataset. Experiments demonstrate that the proposed method significantly outperforms state-of-the-art approaches in both alignment accuracy and visual quality, particularly exhibiting superior robustness and consistency in challenging scenarios.