Score
Designs and implements mechanisms that condition model representations, state updates, or inference-time predictions on geometric information—typically partial or iterative transform signals such as rotations, translations, or orthogonality constraints. Builds modules that inject geometry-aware context into update rules or state-space components so predictions and updates remain consistent with specified geometric constraints.
This work investigates the emergent computational structures in Transformers performing next-token prediction and their explanatory mechanisms for representational geometric features. Method: We propose a theoretical framework of “architecture-constrained parallel Bayesian belief updating,” unifying optimal prediction principles with mechanistic interpretability. Leveraging hidden Markov model (HMM) construction, probability simplex analysis, attention inverse modeling, and constraint-based refinement of optimal prediction equations, we quantitatively predict attention distributions, OV-circuit vector orientations, and embedding manifold geometry. Contribution/Results: Our framework rigorously derives the geometric structure of attention patterns, OV-circuit vectors, and token embeddings, establishing their formal correspondence to Bayesian inference. On controlled HMM tasks, it successfully reproduces and explains canonical geometric representations—including cyclic dynamics and low-dimensional manifolds—demonstrating both quantitative accuracy and mechanistic interpretability of the theoretical predictions.
This work addresses the coverage failure of conformal prediction (CP) under geometric distribution shifts—such as rotations and reflections—where standard CP guarantees degrade. We propose a pose-normalization-augmented CP framework, whose core innovation is the first integration of pose normalization as a geometry-aware feature extractor within the CP pipeline. Crucially, it requires no modification to the underlying black-box predictor and uniformly handles both discrete and continuous geometric transformations while preserving rigorous marginal coverage guarantees. By modeling geometric invariance through normalized features and adapting CP via a standardized interface, our method enhances robustness without compromising CP’s formal statistical assurances. Experiments demonstrate stable empirical coverage ≥95% across diverse geometric shifts, significantly outperforming equivariant models and data-augmentation baselines. Moreover, the framework is fully compatible with arbitrary pre-trained predictors.
This work addresses the limitation of existing state space models in multivariate time series forecasting, which often overlook the dynamically evolving geometric relationships among variables. To remedy this, the paper introduces, for the first time, a symmetric positive definite (SPD) manifold constraint into state space modeling, leveraging Riemannian geometric features as structural regularizers to enable geometry-aware temporal modeling. The proposed method integrates projection from the SPD manifold to its tangent space, a geometry-guided gating mechanism based on manifold-valued signals, and Mamba’s linear-complexity parallel scan architecture. Extensive experiments on eleven real-world benchmark datasets demonstrate state-of-the-art performance, underscoring the critical role of geometric constraints in enhancing forecasting accuracy.
Existing linear probes struggle to uncover the internal encoding structure of geometric information in self-supervised vision Transformers (ViTs). This work proposes a controlled subspace intervention framework that leverages singular value decomposition (SVD) on converged linear probe weights to isolate a low-rank subspace carrying explicit geometric signals. For the first time, subspace analysis reveals distinct differences in geometric representation between DINOv2 and MAE, demonstrating that geometric information is highly compressible, peaks in accuracy at intermediate network layers, and exhibits pronounced low-rank characteristics. These findings provide both theoretical grounding and practical design guidance for lightweight decoders and efficient feature selection strategies in self-supervised vision models.
This work addresses the insufficient integration of learning-based methods and geometric constraints in camera pose and scene structure estimation by proposing a modular framework. The approach first employs a learning model (VGGT) to generate initial hypotheses for depth and relative pose, which are subsequently refined and validated using classical geometric algorithms such as point-to-plane RGB-D ICP. Crucially, the framework explicitly distinguishes the roles of learning as a “proposer” and geometry as a “referee,” emphasizing that the geometric module serves not merely as post-processing but as an essential mechanism for verifying and integrating learned outputs. Experiments on the TUM RGB-D dataset demonstrate that, in moderately challenging rigid scenes, the system significantly outperforms both purely learning-based and purely geometric baselines when the learned depth aligns geometrically with the camera intrinsics and undergoes optimization by the geometric backend.
This work addresses the challenge in generative video editing where object-level geometric manipulations—such as translation, rotation, scaling, duplication, or deletion—often fail to consistently update secondary visual effects like shadows and reflections. To this end, the authors propose GIVE, a unified framework that models pre- and post-edit 3D geometric changes through a consistent object state representation. GIVE employs a dual geometric stream composed of depth and orientation boxes to generate compact, temporally aligned editing instructions. The framework leverages a scalable, procedural synthetic data pipeline built upon a graphics engine for supervised training. GIVE is the first to support diverse geometric editing operations within a single architecture while explicitly modeling 3D state transitions, thereby ensuring consistency in secondary effects, high visual fidelity, temporal coherence, and strong generalization to real-world videos.
This work addresses the lack of a unified modular framework for analyzing adaptive optimizers, which hinders a precise characterization of their behavior under constraints on directional reachability, information budgets, and update rules. We propose a geometric–non-geometric decoupled calculus for optimizers: the geometric module, constituted by a family of positive-definite cometrics, captures realizable descent directions, while the non-geometric module governs mechanisms such as information processing, memory, and control. Within this framework, we establish a direction expressivity theorem and a residual theory for constrained cometric families, disentangling directional expressiveness from condition-number complexity and recasting optimizer design as a Pareto optimization problem under modular budgets. Theoretically, we prove that fully positive-definite geometry exactly spans all strictly descending directions; experiments demonstrate that high-information full-metric probes attain numerical precision on deterministic quadratic problems, and a Muon-style implementation preliminarily validates the auditability of matrix-operator updates.
This work investigates the geometric structure of chain-of-thought reasoning trajectories in large language models and its relationship to task difficulty and answer correctness. Modeling the reasoning process as a discrete curve in the hidden state space, the study introduces an effective dimensionality measure, $d_\rho$, to quantify trajectory complexity and employs spectral analysis alongside geometric functionals derived from positional and kinematic properties to characterize these trajectories. The findings reveal that flatter eigenvalue spectra correspond to more difficult tasks, and that dynamical features extracted from merely the first 20% of generated tokens suffice to effectively predict final answer correctness. Evaluated on the MATH500 dataset, $d_\rho$ achieves an AUC of 0.93 in distinguishing between easy and hard problems, demonstrating both the efficacy and predictive power of the proposed approach.