Score
Designs and trains projection mappings and prototype representations that map embeddings onto learned prototype vectors constrained to lie on a unit hypersphere, and implements prototype-based projection architectures or loss terms that use those hyperspherical anchors. Analyzes and enforces prototype properties (for example, domain-specific or domain-invariant constraints) through objective and constraint design to structure the latent space and support multi-domain generalization.
Existing hyperspherical prototype learning (HPL) methods lack theoretical foundations and are constrained to fixed dimensions, impeding simultaneous geometric controllability and scale invariance. Method: We propose the first principled HPL optimization framework, rigorously proving its global optimality. To overcome dimensional limitations, we construct highly separable class prototypes on arbitrary-dimensional unit hyperspheres via linear group codes, integrating spherical coding theory, convex optimization, and hyperspherical geometric modeling. Contribution/Results: Our framework establishes complete theoretical characterizations—both achievability and converse bounds—for prototype separation. The resulting prototype layouts are provably near-optimal, significantly enhancing inter-class separation and classification robustness across diverse dimensions. Empirical results align closely with theoretical guarantees, demonstrating consistent improvements in both synthetic and real-world benchmarks.
To address the challenge of rigorously enforcing output constraints in safety-critical machine learning applications, this paper proposes a novel hyperspherical constraint representation method. It maps model outputs into a hyperspherical coordinate space centered at the feasible region, thereby intrinsically guaranteeing constraint satisfaction at the representation level—without penalty terms, custom architectures, or post-hoc projection. The approach uniformly supports both convex bounded and star-shaped feasible sets and provides theoretical guarantees of zero constraint violation. Key technical components include a hyperspherical coordinate transformation, geometry-driven feasible set modeling, constraint-aware feature mapping, and a lightweight inverse transformation. Experiments on synthetic and real-world datasets demonstrate that the method achieves prediction accuracy competitive with state-of-the-art constrained learning approaches, incurs no optimization overhead during inference, and strictly maintains zero constraint violations.
This study addresses the limitation of existing abductive explanation methods confined to Euclidean spaces, which hinders their adaptation to non-Euclidean prototypical networks. We extend the abductive latent explanation framework to non-Euclidean geometric manifolds by deriving geometric variant mappings and constructing specialized boundary algorithms to compute subset-minimal formal explanations. This work represents the first extension of formal explanations into non-Euclidean spaces, establishing a unified comparison framework for cross-architecture interpretability. Furthermore, we validate these theoretical constructions on image classification tasks, achieving the first rigorous comparative analysis of interpretability across different architectures. Collectively, this research bridges a critical gap in explainable AI by generalizing formal abductive reasoning beyond Euclidean constraints.
该研究通过引入HGA方法,利用潜空间的几何特征而非配对数据来优化不同潜空间之间的转换,解决了无监督情况下潜空间对齐的问题。
This work investigates low-dimensional-to-high-dimensional (LDHD) generalization—a specialized out-of-distribution generalization problem where training data lies on a low-dimensional submanifold of the high-dimensional test space, equivalent in essence to length generalization. Theoretically, we prove that LDHD generalization is impossible without appropriate inductive bias; min-degree interpolation serves as a universal mechanism for SGD convergence; and chain-of-thought (CoT) reasoning improves length generalization via enhanced positional awareness. Building on these insights, we propose RPE-Square, a novel relative positional encoding that jointly models latent-space scalability and input-format robustness—marking the first unified approach to both properties. Grounded in Boolean function analysis, inductive bias modeling, and optimization dynamics characterization, empirical evaluation demonstrates that RPE-Square significantly outperforms standard RPE on long-sequence tasks. Our framework provides both a unifying theoretical foundation and a practical solution for LDHD and length generalization.
Existing point prototypes in normalized embedding spaces suffer from redundancy or instability due to high intra-class variability of semantic parts, degrading both explanation quality and robustness. This work proposes vMFProto, the first framework to incorporate spherical von Mises-Fisher (vMF) distributions into interpretable classification by modeling each class as a mixture of distributions on the hypersphere, enabling prototypes to adaptively learn concentration parameters that capture part-specific variability. Structured assignment of image patches to prototypes is achieved via entropy-regularized optimal transport, while a distribution-aware diversity regularizer enhances prototype discriminability. A two-stage training strategy facilitates effective prototype discovery followed by end-to-end optimization. Evaluated on CUB-200-2011, Stanford Dogs, and Cars, the method achieves state-of-the-art performance in explanation consistency, stability, and distinctiveness, while maintaining competitive classification accuracy.
研究了在树结构原型网络中,使用双曲几何而非欧氏几何作为潜在流形,能否更好地保持层次分类模型的局部与全局结构。
This study addresses the insufficient discriminability of general-purpose text embeddings in principled evaluation caused by semantic overlap. To mitigate this, we propose Prototype-Guided Contrastive Learning (PGCL), which integrates prototype-anchored attention with shifted margin regularization to generate compact, task-adaptive representations while keeping the encoder frozen. This approach effectively resolves confusion between semantically similar samples with distinct task labels. Experimental results demonstrate that PGCL outperforms original frozen embeddings across three benchmark datasets, achieving particularly significant improvements on Amazon Reviews and matching strong baselines on other tasks. These findings validate the effectiveness of PGCL for parameter-efficient adaptation, offering a robust solution for enhancing embedding distinctiveness without extensive model retraining.
This study addresses the issue in generative 3D optimization where prior distributions overlook sketch styles, leading to design deviations. We propose a condition-free training method based on a rectified flow framework. Exploiting the submanifold properties of the joint distribution, shape latents are optimized via gradient descent during inference, while a token-level cosine penalty is introduced to constrain sketch latents. Furthermore, differentiable drag proxies are incorporated to achieve anchored control. The proposed approach supports design-preserving optimization, explicit dimensional editing, and pure sketch-based synthesis. By ensuring geometric validity while maintaining design fidelity, this work significantly enhances the controllability of customized 3D generation.
This study addresses the limitation of fixed geometric structures in contrastive learning, which struggle to accommodate context-dependent semantic similarity. To overcome this, we propose anchor divergence, a method that integrates contrastive learning, exponential family theory, and information geometry. By establishing a correspondence between anchor distributions and Bregman geometry, our approach defines a context-aware dynamic semantic geometry over fixed representations. Specifically, it directly controls the geometric structure by modeling anchor distributions, thereby enabling adaptive similarity measurement. Experimental results demonstrate that the proposed method efficiently and accurately characterizes context-dependent semantic similarity, yielding substantial improvements in retrieval performance.