Score
Designs and implements supervision signals and loss functions that align learned feature representations with 3D geometric structure, supervising image, voxel, or point-based features using geometry-aware targets or pretrained 3D feature extractors. Builds training pipelines and evaluation criteria that enforce geometric fidelity and structural consistency by imposing correspondence, consistency, or perception-based geometry losses across views or 3D representations.
This paper addresses the fragmented and unsystematic modeling of geometric constraints in deep learning by proposing the first unified taxonomy of geometric constraints tailored for modern deep learning frameworks. Methodologically, it systematically integrates multi-view geometry, epipolar constraints, camera calibration models, self-supervised geometric consistency losses, and differentiable rendering to establish a three-dimensional classification framework spanning modeling principles, integration strategies, and optimization objectives. The contributions are threefold: (1) clarifying the applicability boundaries and failure mechanisms of over one hundred geometric constraints across vision tasks such as depth estimation; (2) uncovering key design paradigms for synergistic co-design of geometric priors and neural architectures; and (3) identifying principled pathways to overcome three core challenges—dynamic scenes, textureless regions, and cross-domain generalization.
Existing linear probes struggle to uncover the internal encoding structure of geometric information in self-supervised vision Transformers (ViTs). This work proposes a controlled subspace intervention framework that leverages singular value decomposition (SVD) on converged linear probe weights to isolate a low-rank subspace carrying explicit geometric signals. For the first time, subspace analysis reveals distinct differences in geometric representation between DINOv2 and MAE, demonstrating that geometric information is highly compressible, peaks in accuracy at intermediate network layers, and exhibits pronounced low-rank characteristics. These findings provide both theoretical grounding and practical design guidance for lightweight decoders and efficient feature selection strategies in self-supervised vision models.
Existing 3D representation learning methods often rely on extrinsic geometry or high-level semantics, making it difficult to capture the intrinsic structure and manifold topology of shapes. This work proposes PRISM, a novel pretraining paradigm that, for the first time, leverages geodesic distance recovery as a self-supervised signal to learn intrinsic geometry through isometric embeddings. To address the inherent imbalance in geodesic distance distributions, the approach introduces a topology-preserving latent space constraint and a two-stage training strategy. The method demonstrates high accuracy, robustness, and efficiency in geodesic distance prediction and achieves state-of-the-art performance on downstream tasks including shape recognition, surface parameterization, and non-rigid correspondence.
This work addresses the challenges of few-shot known-class detection, unknown-class rejection, and continual learning of novel classes in open-world object detection. To this end, the authors propose a Category Geometry Supervision (CGS) framework, which, for the first time, leverages the geometric structure of inter-class relationships as a supervisory signal in detection models. By introducing a category geometry alignment loss in the prototype representation space, CGS preserves the visual-semantic dissimilarity structure among categories estimated from training data. Integrated within a prototype-based detection architecture alongside standard task losses, the method substantially improves performance across few-shot detection, open-set recognition, and incremental novel-class learning scenarios. Experiments on benchmarks such as COCO demonstrate consistent gains in few-shot detection accuracy, novel-class incorporation, and unknown-class recall, without compromising known-class detection performance.
This work addresses the problem of open-domain single-image 3D geometric reconstruction. To resolve global scale and translation ambiguities, we propose an affine-invariant 3D point cloud representation. Methodologically, we design an optimal point cloud alignment solver and a multi-scale local geometric consistency loss to mitigate the inherent ambiguity of monocular geometric supervision. Our approach integrates affine-invariant representation learning, robust point cloud registration, and end-to-end training on a hybrid large-scale dataset. Extensive experiments demonstrate that our method achieves state-of-the-art performance across multiple unseen benchmarks. It significantly improves accuracy and generalization in monocular 3D point cloud reconstruction, depth estimation, and field-of-view prediction. By eliminating the need for camera calibration or explicit metric priors, our framework establishes a new paradigm for uncalibrated single-image geometric understanding.
This work investigates how training dynamically shapes the Riemannian geometric structure induced by neural network representations. **Problem**: While deep networks operate in high-dimensional feature spaces, the geometric evolution of their induced Riemannian metric during training remains poorly understood. **Method**: We integrate Riemannian geometric analysis, infinite-width network theory, and feature-space metric modeling, conducting systematic experiments across supervised and self-supervised learning paradigms. **Contribution/Results**: We theoretically prove that infinitely wide random networks initially possess an isotropic (highly symmetric) Riemannian metric; training actively breaks this symmetry by locally amplifying the metric tensor near decision boundaries—a phenomenon we term *boundary-sensitive metric amplification*. Empirical validation across deep image classification and self-supervised learning confirms its robustness, revealing a self-emergent geometric inductive bias. This provides a novel geometric paradigm for understanding the intrinsic geometry of nonlinear feature learning.
The field of 3D vision suffers from fragmented data representations, learning paradigms, and benchmarking protocols, leading to a lack of unified understanding regarding efficiency, fidelity, and scalability. This work proposes the first cohesive conceptual framework that integrates geometric representations—such as point clouds, meshes, voxels, and 3D Gaussians—with diverse learning paradigms—including 2D-supervised learning, implicit neural representations, and 4D modeling—and connects them to real-world application scenarios. By constructing a structured knowledge graph of 3D vision, the study systematically relates dataset design, supervision mechanisms, and task requirements, clarifying the trade-offs between efficiency and fidelity and charting pathways for multimodal geometric grounding. This framework offers systematic guidance for reconstruction, generation, and dynamic scene modeling, advancing the field toward a unified and efficient paradigm.
This work addresses the fundamental challenge that 2D vision models cannot be directly applied to irregular and sparse 3D data such as point clouds and meshes. To this end, it proposes the first unified taxonomy encompassing data representations, architectural designs, and hybrid strategies. Existing approaches are systematically categorized into three paradigms: projection-based data-centric methods, architecture-centric techniques leveraging native 3D structures, and hybrid approaches combining both. The study provides a thorough analysis of the trade-offs among computational complexity, reliance on pretraining, and preservation of geometric inductive biases. By integrating pathways including transfer from 2D CNNs or Vision Transformers, native 3D network design, multimodal fusion, and self-supervised learning, this work offers a systematic roadmap for advancing 3D foundation models, geometric self-supervised learning, and multimodal representation learning.
Existing CAD learning approaches discretize B-Rep models into triangle meshes, thereby discarding the analytical surface representations and topological information essential for consistent instance-level analysis. This work proposes STEP-Parts, a deterministic pipeline that directly extracts geometric instance partitions from native STEP B-Rep data. The method defines partitions based on intrinsic B-Rep topology, merges faces using analytical surface types and near-tangent plane continuity criteria, and transfers labels to triangulated meshes via face-to-mesh correspondence mapping. STEP-Parts ensures boundary consistency across varying triangulations and processes the DeepCAD subset of the ABC dataset—comprising approximately 180,000 models—in under six hours. The resulting labels significantly enhance performance in implicit reconstruction-segmentation tasks and point cloud networks. Code and precomputed labels are publicly released.
This work addresses the geometric hallucinations commonly observed in existing Point-Vision-Language Models during 3D structure generation, which stem from sparse geometric signals being overwhelmed by global rewards in reinforcement learning. To mitigate this issue, the authors propose a geometric reward credit assignment mechanism that decouples holistic supervision into domain-specific signals and accurately propagates them to the corresponding tokens. Additionally, they introduce a reprojection consistency constraint as a cross-modal verifier to ensure the physical plausibility of 3D predictions. Evaluated on the ShapeNetCore benchmark, the method achieves significant performance gains, attaining a 3D keypoint accuracy (KPA) of 0.93, a 3D bounding box IoU of 0.686, and a reprojection consistency of 0.852, while preserving strong 2D localization capabilities.
Monocular endoscopic visual navigation faces significant challenges in pose estimation, depth prediction, and image-to-anatomy alignment due to scarce depth cues, weak tissue texture, non-rigid deformations, and cross-domain appearance variations. This work proposes a unified framework that leverages synthetic data to provide geometric supervision and introduces hierarchical perception-aware low-rank adapters—replacing standard LoRA—embedded at multiple layers of a vision foundation model. By integrating layered geometric-semantic training objectives, the method jointly optimizes geometric correspondence in intermediate features and semantic consistency in deep representations. The approach substantially enhances both geometric fidelity and semantic quality of cross-domain features, improving pose and monocular depth estimation on both public and private datasets. It enables effective transfer from synthetic bronchoscopic data to real-world scenes and supports rapid few-shot adaptation to new domains such as sinus and colonoscopy.