Score
Designs and evaluates methods that estimate the low‑dimensional manifold structure of a representation space and construct projection operators that map update or intervention vectors onto that manifold to constrain edits or updates. Builds and analyzes manifold estimators, projection routines, and intervention protocols that restrict changes to manifold subspaces so as to preserve non‑target information while enabling targeted representation modifications.
Traditional linear dimensionality reduction methods often fail to effectively uncover the intrinsic low-dimensional manifold structure embedded in high-dimensional data. This work systematically traces the historical development of manifold fitting and, for the first time, categorizes it into three distinct phases: nonparametric statistics, mathematically inspired analysis, and modern practical statistics. It clarifies manifold fitting’s role as an independent geometric data analysis tool and delineates its conceptual boundaries from related techniques such as manifold embedding and denoising. By integrating nonparametric methods, differential geometry, and contemporary statistical learning approaches, the paper explores cutting-edge applications of manifold fitting in neural networks and bioinformatics, offering a comprehensive reference framework that elucidates both its theoretical limits and practical utility.
Existing anomaly detection methods typically assume that normal data occupy a non-zero volume in the ambient space, overlooking their intrinsic geometric structure as lying on a low-dimensional manifold, which limits performance. This work proposes a novel manifold projection paradigm: learning a projection operator that maps inputs onto the manifold of normal samples and using the projection residual as the anomaly criterion. By avoiding explicit modeling of the degenerate data distribution, the approach prevents misclassifying rare yet normal instances and provides a unified explanation for both the effectiveness and failure modes of reconstruction-based methods. Extensive experiments demonstrate that the proposed framework significantly outperforms conventional boundary-learning approaches and achieves state-of-the-art results across multiple benchmarks compared to existing reconstruction-based models.
Existing inverse projection methods are constrained by fixed manifold structures, limiting their ability to adequately capture the diversity of high-dimensional image data and hindering their applicability in tasks such as data augmentation, classifier analysis, and data imputation. This work proposes a general and controllable inverse projection framework that enables flexible exploration and reconstruction of the high-dimensional space underlying any dimensionality reduction technique—such as t-SNE or UMAP—through two intuitive, user-defined parameters. By transcending conventional structural limitations, the method offers broad compatibility, ease of implementation, and strong user-guided control. Its effectiveness is demonstrated in applications like image style transfer, where it achieves more comprehensive and practical coverage of the high-dimensional data space.
High-dimensional data often reside on low-dimensional manifolds, yet existing manifold dimension estimation algorithms lack systematic evaluation and reproducible benchmarks. Method: We conduct a comprehensive empirical assessment of eight representative methods across synthetic and real-world datasets, quantifying the effects of noise, curvature, and sample size on estimation accuracy. We introduce a dataset-aware hyperparameter tuning principle and establish a controlled-variable experimental framework integrating local linear embedding, nearest-neighbor statistics, and multiscale geometric analysis. Contribution/Results: Contrary to the “complexity implies superiority” assumption, simple methods—such as the nearest-neighbor distance ratio and PCA-based gradient estimation—consistently achieve higher accuracy and robustness across most scenarios. This work delivers the first open-source, fully reproducible benchmark for manifold dimension estimation, accompanied by practical guidelines. It provides both theoretical insight and empirical evidence to inform unsupervised learning and dimensionality reduction methodology selection.
This paper addresses the out-of-sample embedding problem: efficiently and robustly embedding new samples into an existing vector space given similarity or dissimilarity data. We propose a unified theoretical framework that, for the first time, systematically categorizes existing kernel-based methods into two paradigms—kernel-based projection and constrained reconstruction—and rigorously establish their mathematical equivalence and applicability boundaries. Our analysis reveals that constrained reconstruction reduces to a unidimensional search and elucidates its statistical robustness mechanism. Leveraging these insights, we design a computationally efficient, noise-resilient constrained reconstruction algorithm. Experiments demonstrate that our method significantly outperforms state-of-the-art approaches under high noise and sparse neighborhood conditions. The proposed framework provides a principled foundation for incremental updates and practical deployment of graph embeddings.
Diffusion models in long-horizon, sparse-reward offline reinforcement learning often deviate from the data manifold due to inaccurate guidance, generating infeasible trajectories—limiting their deployment in safety-critical applications. To address this, we propose LoMAP (Local Manifold Approximation and Projection), a training-free method that constructs a low-rank manifold subspace from offline data and applies real-time correction to diffusion-guided samples via local PCA and orthogonal projection, ensuring trajectory feasibility. We establish, for the first time, a theoretical lower bound linking the guidance gap to manifold deviation, thereby providing formal feasibility guarantees for diffusion-based planning. LoMAP is modular and seamlessly integrates with existing hierarchical diffusion planners. Empirical evaluation on standard offline RL benchmarks demonstrates significant improvements in both trajectory feasibility and task success rates.
This study addresses the challenging problem of fitting an unknown number of hyperplanes to data, which involves non-convexity, non-differentiability, and uncertainty in model order. To tackle these difficulties, the authors propose a two-stage unsupervised learning approach grounded in a unit-sphere manifold framework. In the first stage, they integrate Riemannian expectation-maximization with heavy-tailed kernel density estimation to robustly infer posterior probabilities. The second stage employs hard assignment annealing to obtain geometrically consistent local optima. Key innovations include a manifold optimization framework for handling non-convex constraints, a projection-based density estimation scheme for initialization, and the two-stage optimization strategy itself. Experimental results demonstrate that the proposed method significantly outperforms state-of-the-art baselines in both geometric accuracy and robustness.
This work addresses the challenge of achieving strict idempotence in generative models under repeated application, where output drift arises due to geometric inconsistencies between the data manifolds learned by the encoder and decoder. The study identifies this manifold misalignment as the key cause of idempotence failure—a previously unexamined issue—and introduces a novel training framework that explicitly aligns the geometric structures of both components. By enforcing the encoder’s projection and the decoder’s reconstruction to share a common underlying manifold during training, the proposed method substantially reduces idempotence error, yielding perfectly consistent outputs across repeated generations. Empirical results demonstrate significant improvements in identity preservation and information stability for image generation and editing tasks.
This work addresses how to uncover interpretable concept manifolds embedded within the stacked representations of language models. The authors propose Manifold Probe, a method that generalizes traditional linear probing to manifold probing by integrating supervised manifold learning with linear predictability analysis. This approach identifies continuous geometric structures in representation space corresponding to high-level concepts—such as time or space—and determines their encoding directions. Beyond merely detecting the presence of such concepts, the method enables causal intervention: manipulating activations along discovered manifold directions directly alters model behavior. Experiments on Llama 2-7B demonstrate that perturbing representations along the extracted temporal manifold significantly shifts the model’s generated outputs regarding the release years of cultural works, thereby validating both the interpretability and causal efficacy of the recovered manifolds.
This work addresses the challenge of precisely controlling specific behaviors—such as refusal or sycophancy—in large language models, where targeted interventions often produce unintended side effects. The authors propose a low-rank subspace diagnostic framework that reveals, for the first time, that distinct behaviors share internal representations in activation space. Through geometric analysis of decision subspaces and the mean squared cosine of principal angles, they demonstrate that intervention effects propagate asymmetrically, depending on the degree of subspace overlap and the angular proximity to the decision subspace. Experiments across multiple instruction-tuned models (7B–70B) show that behaviors exhibiting high representational overlap and closer alignment with the decision subspace are more susceptible to intervention, thereby explaining the fundamental difficulty in achieving independent behavioral control.
This work addresses optimization problems defined over products of simplices, such as low-rank learning of discrete multivariate probability distributions and function data registration based on the Square-Root Velocity Function (SRVF) representation. To tackle the inherent constraints, the authors propose a smooth reparameterization that is strictly convex element-wise, transforming the constrained problem into an unconstrained optimization over a Riemannian manifold. The resulting problem is solved via Riemannian gradient descent (RGD). Theoretical analysis shows that this reparameterization maps second-order KKT points on the manifold to weak second-order KKT points of the original problem, ensuring theoretical soundness while enhancing computational efficiency. Experiments demonstrate that RGD significantly outperforms projected gradient descent (PGD), achieving more accurate shape-preserving registration in functional data and efficiently solving probability tensor decomposition tasks.