Score
Designs and implements representations and algorithms that construct explicit, parametric ray-centric geometric features (e.g., origin, direction, sampled depths and along-ray encodings), encode spatial relationships and orientation information along those rays, and provide methods to aggregate or fuse these ray-based features with other feature representations to support downstream geometric reasoning or perception tasks.
This work addresses the insufficient integration of learning-based methods and geometric constraints in camera pose and scene structure estimation by proposing a modular framework. The approach first employs a learning model (VGGT) to generate initial hypotheses for depth and relative pose, which are subsequently refined and validated using classical geometric algorithms such as point-to-plane RGB-D ICP. Crucially, the framework explicitly distinguishes the roles of learning as a “proposer” and geometry as a “referee,” emphasizing that the geometric module serves not merely as post-processing but as an essential mechanism for verifying and integrating learned outputs. Experiments on the TUM RGB-D dataset demonstrate that, in moderately challenging rigid scenes, the system significantly outperforms both purely learning-based and purely geometric baselines when the learned depth aligns geometrically with the camera intrinsics and undergoes optimization by the geometric backend.
Existing linear probes struggle to uncover the internal encoding structure of geometric information in self-supervised vision Transformers (ViTs). This work proposes a controlled subspace intervention framework that leverages singular value decomposition (SVD) on converged linear probe weights to isolate a low-rank subspace carrying explicit geometric signals. For the first time, subspace analysis reveals distinct differences in geometric representation between DINOv2 and MAE, demonstrating that geometric information is highly compressible, peaks in accuracy at intermediate network layers, and exhibits pronounced low-rank characteristics. These findings provide both theoretical grounding and practical design guidance for lightweight decoders and efficient feature selection strategies in self-supervised vision models.
This paper presents a systematic review of learning-based 3D representations for tasks such as 3D reconstruction, novel view synthesis, and rendering, spanning from traditional explicit formulations—including meshes, point clouds, and voxels—to emerging implicit neural fields and primitive-based splatting techniques like 3D Gaussian Splatting. Emphasizing the evolutionary trajectory of 3D representations themselves, the work highlights the paradigm shift from explicit to implicit modeling, clarifying the mathematical formulations, strengths, limitations, and suitable application scenarios of each approach. Distinct from prior task-centric surveys, this study centers on representation as the core organizing principle, offering fresh insights for 3D/4D content generation and identifying key challenges and future research directions to serve as a theoretical reference for the computer graphics and vision communities.
This work addresses the ambiguity inherent in traditional scalar representations of angular data, which fail to distinguish between nearby angles differing by more than π due to periodicity. To resolve this, the authors propose a high-dimensional, real-valued distributed representation based on Fourier embeddings, integrated with spatial semantic pointers to enable neurally interpretable encoding of periodic signals. They formalize the Dirichlet kernel and the periodic Gaussian kernel within this framework, allowing flexible and theoretically grounded control over angular similarity measures. The resulting approach provides unambiguous representations for arbitrarily close angles and establishes a principled design framework for similarity functions with customizable kernel shapes and provable theoretical guarantees.
Existing vision-language models struggle to effectively leverage geometric information for spatial reasoning in both static and dynamic scenes. To address this limitation, this work proposes GeoSR, a framework that weakens 2D visual shortcuts through geometry-aware masking and adaptively enhances the contribution of geometric tokens in critical regions via a gated routing mechanism. GeoSR integrates geometric tokens—generated by a pretrained 3D foundation model—into the vision-language model using a masked integration strategy and gated fusion. Experimental results demonstrate that GeoSR achieves state-of-the-art performance across multiple benchmarks for spatial reasoning in both static and dynamic settings, significantly improving the model’s geometric perception and reasoning capabilities.
This study addresses the challenge of constructing rotation-invariant vector representations for planar shapes by proposing a method that strictly encodes star-shaped normalized contours into Euclidean vectors. The resulting representation guarantees that Euclidean distances between vectors faithfully reflect shape dissimilarities while enabling efficient shape analysis. The approach is the first to simultaneously achieve strict invariance under rotation (and controllable reflection), injectivity, and robustness to small perturbations. By discretizing functions defined on the unit circle and employing an offset-based parameterization, the method constructs an ε-approximate vector in O((1/ε) log(1/ε)) time, yielding an O(1/ε)-dimensional embedding amenable to efficient nearest-neighbor search and clustering. Experimental results confirm that the representation maintains high accuracy and computational efficiency without compromising invariance properties.
This work addresses the challenge of simulating human-like multi-step logical reasoning with auxiliary constructions in geometric problem solving by proposing a novel framework that integrates mathematical reasoning with procedural representations. The approach employs program code as an intermediate visual representation, decoupling discovery reasoning from code generation in a latent space and structuring the reasoning manifold through supervised fine-tuning. The study demonstrates that hierarchical syntactic code structures effectively encode rich mathematical semantics, offering greater expressiveness than purely visual representations. Experimental results show that the proposed method significantly enhances geometric reasoning performance while yielding clearer and more interpretable multi-step derivations.
Existing video generation models suffer from poor cross-frame 3D consistency and limited camera controllability due to the absence of explicit 3D structural modeling. To address this, this work proposes RayPE—a novel ray-space positional encoding based on Plücker coordinates—that is, for the first time, additively integrated into the self-attention mechanism of video diffusion Transformers. This integration naturally decomposes attention scores into content, geometry, and their interaction terms. Combined with QK-swapped attention, gating mechanisms, and synergistic normalization using RMSNorm and QKNorm, RayPE achieves substantial improvements in camera controllability, 3D consistency, and overall video quality on mixed multi-view camera data, while introducing less than 0.1% additional parameters.