Score
Designs and analyzes mathematical distance and similarity measures and their associated scale-space representations (e.g., kernels, filter banks, or feature transforms) that are invariant under rotations. This work includes constructing rotationally‑invariant scale‑space variants, proving stability to angular perturbations, and implementing metrics for robust orientation‑independent comparisons.
This study addresses the challenge of constructing rotation-invariant vector representations for planar shapes by proposing a method that strictly encodes star-shaped normalized contours into Euclidean vectors. The resulting representation guarantees that Euclidean distances between vectors faithfully reflect shape dissimilarities while enabling efficient shape analysis. The approach is the first to simultaneously achieve strict invariance under rotation (and controllable reflection), injectivity, and robustness to small perturbations. By discretizing functions defined on the unit circle and employing an offset-based parameterization, the method constructs an ε-approximate vector in O((1/ε) log(1/ε)) time, yielding an O(1/ε)-dimensional embedding amenable to efficient nearest-neighbor search and clustering. Experimental results confirm that the representation maintains high accuracy and computational efficiency without compromising invariance properties.
This study addresses the stability of image metrics in Gaussian scale space under geometric deformations and additive noise. To overcome the sensitivity of classical scale-space metrics to rotation, the authors propose a rotation-invariant variant and develop a novel metric that is robust to both geometric transformations and noise by integrating tools from harmonic analysis, optimal transport theory, and numerical algorithms. The resulting metric admits efficient computation from finite samples and is theoretically linked to the Wasserstein distance and Besov spaces. Experimental results demonstrate its enhanced stability and effectiveness, significantly improving the reliability of image comparison in perturbed settings.
This study investigates the geometric structure of image manifolds induced by 3D object poses and their inter-class variations, aiming to explain the success of visual representation learning from a differential-geometric perspective. Method: We propose a novel framework integrating geometry-preserving manifold learning with Kendall shape theory: (i) modeling pose-induced image manifolds as smooth, nonlinear manifolds in latent space; and (ii) introducing a rigidity-invariant Kendall metric to enable shape-invariant quantification and clustering of manifold geometry. Contribution/Results: We systematically demonstrate, for the first time, that image manifolds of the same object class exhibit significant clustering in shape space—and crucially, the degree of such clustering correlates with model generalization performance. Our approach enables comparable geometric modeling across object classes, providing both theoretically grounded, geometrically interpretable principles and practical guidance for designing vision algorithms.
This paper addresses the problem of visual relative pose estimation. We propose a novel modeling framework based on dual rotation parameterization: jointly optimizing the rotation matrices of two cameras directly on the SO(3) manifold, bypassing conventional essential matrix decomposition or end-to-end pose regression paradigms. Our method introduces a geometrically grounded coordinate transformation and three differentiable, geometrically consistent energy functions, minimized jointly within a Riemannian optimization framework. The approach achieves strong robustness, high accuracy, and excellent generalization across diverse relative pose tasks—including two-view pose estimation and Structure-from-Motion (SfM) initialization—outperforming state-of-the-art methods by significant margins. To foster reproducibility and community advancement, we release our source code, demonstration videos, and benchmark datasets.
This study addresses the limitations of 3D shape descriptors that are sensitive to parameterization, pose, and scale while failing to distinguish mirror chirality. To overcome these issues, we propose a shape description method that is both complete and stable. The approach eliminates geometric dependencies through conformal mapping and conformal barycentric normalization. Furthermore, by leveraging classical invariant theory to identify harmonic bands as binary forms, it constructs polynomial invariants of the rotation group to precisely decouple transformation dependencies while preserving chirality information. Benchmark evaluations demonstrate that the proposed descriptor exhibits excellent stability and completeness, successfully achieving accurate separation between mirror-symmetric and asymmetric pairs within bilateral anatomical structures.
This work addresses the significant performance degradation of existing deep image matching methods under large in-plane rotations. Through systematic investigation of where to best incorporate rotation invariance within sparse feature matching pipelines, extensive training, and multi-benchmark evaluation, the study demonstrates that introducing rotation invariance solely at the descriptor stage achieves robustness comparable to that of rotation-invariant matchers while being more computationally efficient. Moreover, it shows that, with sufficient training data, rotation invariance does not compromise general matching performance and highlights the critical role of data scale in enabling robust rotation generalization. The released models achieve state-of-the-art results on benchmarks including WxBS, HardMatch, and SatAst, substantially improving matching robustness across multimodal, extreme-viewpoint, and satellite imagery scenarios.
This study addresses the insufficient robustness and accuracy of local feature matching in overlapping regions of satellite imagery. To this end, the authors construct a manually curated satellite image dataset annotated with GPS coordinates and conduct a systematic evaluation of SIFT and ORB algorithms across the entire matching pipeline—including keypoint detection, descriptor extraction, feature matching, and RANSAC-based geometric verification. Using the inlier ratio as the primary metric for matching quality, the work quantitatively analyzes the impact of keypoint quantity on matching performance. The results reveal a nonlinear relationship between the number of detected keypoints and the inlier ratio, offering empirical evidence and theoretical guidance for algorithm selection and parameter tuning in remote sensing image matching tasks.
This work addresses the challenge of rotation-invariant object recognition under data-scarce conditions by proposing a compact deep network architecture that integrates a spectral-spatial polar representation. The method uniquely embeds a spectral-spatial polar structure into a lightweight neural network, mathematically guaranteeing strict rotational invariance of features without relying on data augmentation. Experimental results demonstrate that, in low-data regimes, the proposed model significantly outperforms conventional convolutional neural networks and achieves theoretically provable rotation-invariant recognition performance.
This study addresses the reliance on manual annotation and limited scalability of existing 3D rotational symmetry labeling by proposing a fully automated, reference-free analytical framework. Methodologically, it introduces a novel hierarchy-guided symmetric structure reconstruction strategy that enables classification across eight symmetry types and full-axis localization. Furthermore, a texture-aware mechanism is incorporated to mitigate order degradation caused by object appearance, combined with self-consistency analysis to infer rotational orders. Experiments demonstrate that the proposed framework achieves 94.75% accuracy on the GSO dataset. Additionally, integrating the derived symmetry priors into FoundationPose yields accuracy improvements of up to 0.9% across five BOP benchmarks, validating its practical utility in 6D pose estimation tasks.
This work addresses the challenge of 6D object pose estimation for symmetric objects, which is inherently ambiguous due to rotational symmetries. Existing approaches typically rely on custom loss functions, specialized network architectures, or symmetry-invariant metrics. In contrast, this paper introduces SARR—a novel rotation representation based on a symmetry-aware refinement of trigonometric identities—that resolves symmetry-induced ambiguity directly at the level of rotation representation. Notably, SARR produces visually continuous and unique canonical poses without requiring 3D object models or explicit symmetry priors. It integrates seamlessly with standard CNNs and operates using only depth maps or textureless RGB/gray-scale images as input. Evaluated on the T-LESS and ITODD benchmarks, SARR significantly outperforms state-of-the-art methods under the symmetry-sensitive AR_C metric and also achieves superior performance under conventional symmetry-invariant metrics, surpassing multiple mainstream rotation representations.