Score
Design and implement representations and models that extract and encode geometric and task constraints from images at the patch or region level, producing embeddings or geometry-aware regression outputs that capture relations (e.g., distances, angles, feasibility) between visual elements. These encodings are tailored to different constraint types so downstream planners, constraint solvers, or reasoning modules can operate directly on visual inputs.
This paper addresses the fragmented and unsystematic modeling of geometric constraints in deep learning by proposing the first unified taxonomy of geometric constraints tailored for modern deep learning frameworks. Methodologically, it systematically integrates multi-view geometry, epipolar constraints, camera calibration models, self-supervised geometric consistency losses, and differentiable rendering to establish a three-dimensional classification framework spanning modeling principles, integration strategies, and optimization objectives. The contributions are threefold: (1) clarifying the applicability boundaries and failure mechanisms of over one hundred geometric constraints across vision tasks such as depth estimation; (2) uncovering key design paradigms for synergistic co-design of geometric priors and neural architectures; and (3) identifying principled pathways to overcome three core challenges—dynamic scenes, textureless regions, and cross-domain generalization.
This work addresses the limitation of existing visual representation methods, which overly rely on global geometric structure and struggle to effectively model compositional relationships among elements. Through a systematic evaluation of 21 visual encoders, the study reveals—for the first time—that standard geometric metrics are nearly uncorrelated with compositional binding capacity. To address this gap, the authors propose functional sensitivity, measured via the input–output Jacobian matrix, as a complementary evaluation dimension. Integrating geometric statistics, Jacobian analysis, and theoretical derivation, they demonstrate that functional sensitivity reliably predicts compositional binding performance and elucidate its origin at the level of optimization objectives. This insight establishes a novel evaluation paradigm for representation learning that moves beyond conventional geometric assessments.
This work investigates the geometric reasoning capabilities of Graph Neural Networks (GNNs) and Transformers in embedding space, focusing on reconstructing implicit 2D geometric structures—specifically, predicting spatial coordinates and recovering underlying shapes from point sets defined by discrete geometric constraints on a 2D grid. Method: We propose a geometry-aware GNN architecture explicitly designed for geometric reasoning. Crucially, it operates without explicit coordinate supervision, relying solely on relational graph structure. Contribution/Results: We demonstrate, for the first time, that the learned node embeddings spontaneously organize into a low-dimensional subspace preserving neighborhood relationships—effectively recovering the latent 2D grid topology. Quantitatively, our GNN significantly outperforms Transformer baselines in both prediction accuracy and scalability. Qualitative analysis confirms that the embedding space faithfully encodes geometric structure, providing strong evidence of implicit geometric modeling capacity. These findings establish a novel, interpretable paradigm for spatial reasoning grounded in learned embeddings.
Existing linear probes struggle to uncover the internal encoding structure of geometric information in self-supervised vision Transformers (ViTs). This work proposes a controlled subspace intervention framework that leverages singular value decomposition (SVD) on converged linear probe weights to isolate a low-rank subspace carrying explicit geometric signals. For the first time, subspace analysis reveals distinct differences in geometric representation between DINOv2 and MAE, demonstrating that geometric information is highly compressible, peaks in accuracy at intermediate network layers, and exhibits pronounced low-rank characteristics. These findings provide both theoretical grounding and practical design guidance for lightweight decoders and efficient feature selection strategies in self-supervised vision models.
Existing RGB-based imitation learning methods rely on generic vision encoders (e.g., ResNet, ViT) that lack explicit 3D geometric modeling capabilities, limiting robotic manipulation performance. To address this, we propose eVGGT—a lightweight, geometry-aware visual encoder distilled from VGGT—retaining strong 3D reasoning while achieving an 8.7× inference speedup and 80% parameter reduction. eVGGT seamlessly integrates into mainstream imitation learning frameworks (e.g., ACT, DP), jointly processing RGB inputs and explicit geometric representations for policy learning. Evaluated in both simulation and real-robot settings, eVGGT improves single- and dual-arm manipulation success rates by up to 6.5% over baseline methods. It thus bridges high-fidelity 3D understanding with real-time deployability, advancing geometrically grounded visuomotor control.
This study investigates the true role and underlying mechanism of positional encoding in Vision Transformers (ViTs) with respect to spatial geometric reasoning. Addressing the lack of deep understanding regarding the geometric significance of positional encoding in existing literature, we propose a token-level multi-view geometric consistency diagnostic framework, which for the first time demonstrates that positional encoding acts as a causal factor in shaping the spatial structure of ViT representations. Through comprehensive ablation and probing experiments across 14 foundational ViT models, we validate that positional encoding simultaneously guides both local structural coherence and global layout organization, thereby establishing its essential role as a critical geometric prior in ViT architectures.
This study addresses the lack of systematic analysis regarding the geometric evolution of internal representations during Vision Transformer (ViT) training. The authors propose the TGO-II framework, which integrates Centered Kernel Alignment (CKA), Singular Vector Canonical Correlation Analysis (SVCCA), TwoNN intrinsic dimension estimation, and token covariance analysis. Applying this framework to ViT-Small/16 under supervised training, they uncover a tripartite geometric evolution pattern: progressive layer-wise specialization, an initial rise followed by stabilization of intrinsic dimensionality, and the persistent presence of strong token interaction structures. These findings demonstrate that increased representational complexity co-occurs with layer specialization without requiring token decorrelation, thereby challenging the conventional assumption that complexity arises from token independence and highlighting ViT’s capacity to achieve rich representational transformations through sustained strong token interactions.
This work addresses the fundamental challenge that 2D vision models cannot be directly applied to irregular and sparse 3D data such as point clouds and meshes. To this end, it proposes the first unified taxonomy encompassing data representations, architectural designs, and hybrid strategies. Existing approaches are systematically categorized into three paradigms: projection-based data-centric methods, architecture-centric techniques leveraging native 3D structures, and hybrid approaches combining both. The study provides a thorough analysis of the trade-offs among computational complexity, reliance on pretraining, and preservation of geometric inductive biases. By integrating pathways including transfer from 2D CNNs or Vision Transformers, native 3D network design, multimodal fusion, and self-supervised learning, this work offers a systematic roadmap for advancing 3D foundation models, geometric self-supervised learning, and multimodal representation learning.
Achieving compositional generalization in vision models under unseen compositional scenarios remains a fundamental challenge for foundation models. This work theoretically establishes, for the first time, that effective compositional generalization necessitates representations that are disentangled, transferable, and stable—properties that require conceptual components in the embedding space to be linearly decomposable and mutually orthogonal. We further derive a geometric constraint linking the number of concepts to the embedding dimensionality. Guided by this theory, we conduct empirical analyses on prominent models including CLIP, SigLIP, and DINO, revealing that their representations consistently exhibit partial linear factorization and low-rank near-orthogonal structures. Crucially, the degree of such structure strongly correlates with compositional generalization performance, thereby validating the proposed theoretical framework.
This study addresses the limited geometric understanding of current vision–language–action (VLA) models, which constrains their performance in embodied tasks. For the first time, the authors quantify the “geometry gap” between VLAs and geometric foundation models (GFMs) using linear probing, and systematically evaluate—under a unified experimental setup—the impact of three fusion architectures, training data scale, and multi-view inputs on geometric perception. The findings reveal that specific fusion architectures substantially enhance geometric comprehension, while multi-view observations and sufficient training data are critical for high performance. This work establishes key design principles and provides empirical evidence for developing geometry-aware VLA systems.