Score
Design and implement analyses and diagnostic tools that characterize the geometric structure of learned latent spaces and manifolds, including metrics and visualizations of curvature, connectivity, local linearity, and interpolation paths. Use these analyses to assess how edits or hyperparameter changes shift the latent landscape, evaluate interpolation quality, and identify geometry-induced locality violations and editor-specific failure modes.
Traditional linear dimensionality reduction methods often fail to effectively uncover the intrinsic low-dimensional manifold structure embedded in high-dimensional data. This work systematically traces the historical development of manifold fitting and, for the first time, categorizes it into three distinct phases: nonparametric statistics, mathematically inspired analysis, and modern practical statistics. It clarifies manifold fitting’s role as an independent geometric data analysis tool and delineates its conceptual boundaries from related techniques such as manifold embedding and denoising. By integrating nonparametric methods, differential geometry, and contemporary statistical learning approaches, the paper explores cutting-edge applications of manifold fitting in neural networks and bioinformatics, offering a comprehensive reference framework that elucidates both its theoretical limits and practical utility.
This paper addresses the challenges of evaluating dimensionality reduction (DR) effectiveness and estimating intrinsic data dimensionality. We propose a geometric profiling method based on sectional curvature in discrete metric spaces, which characterizes large-scale data geometry via metric relationships among point triplets. For the first time, this approach systematically introduces differential-geometric curvature into quantitative DR quality assessment and intrinsic dimension estimation—without requiring embedded coordinates or manifold assumptions, thus ensuring both theoretical rigor and computational feasibility. Experiments across diverse synthetic and real-world datasets demonstrate that our method robustly discriminates DR algorithm performance, achieves significantly lower intrinsic dimension estimation error than state-of-the-art methods (e.g., MDS- and PCA-based estimators), and successfully uncovers latent negative curvature in empirical networks—including social and biological networks.
Internet latency data exhibit complex topological structures that are difficult to interpret intuitively. Method: This paper proposes a manifold-based geometric visualization framework that models real-world latency measurements as a geography-aware manifold embedded in a 2D Euclidean space. It employs geodesic distance as the intrinsic metric and—novelty introduced herein—uses Forman–Ricci curvature to quantify local connectivity and detect anomalies. The approach integrates graph neural networks, nonlinear dimensionality reduction (t-SNE/UMAP), and GIS projection to enable curvature-driven interactive rendering. Contribution/Results: Implemented as the Matisse system, the method successfully identifies high-curvature anomalous regions in U.S. public Internet latency data, demonstrating the efficacy of geometric representation for uncovering performance bottlenecks and topological vulnerabilities. Its core innovation lies in incorporating Ricci curvature into latency-space modeling, establishing a tri-coupled visualization paradigm linking latency, geometry, and geography.
This work addresses the lack of a universal statistical interpretation for the manifold hypothesis—that high-dimensional data approximately reside on low-dimensional manifolds. We propose the Latent Metric Model (LMM), a generative framework grounded in fundamental statistical concepts: latent variables, variable dependence, and stationarity—providing the first unified statistical justification for the manifold assumption. Methodologically, LMM integrates neighborhood graph construction, spectral analysis, and an interpretable inference framework to enable unsupervised manifold discovery and geometric structure recovery under weak priors. Experiments demonstrate that complex manifold geometries naturally emerge from minimal statistical mechanisms; LMM significantly reduces reliance on hand-crafted priors on both synthetic and real-world datasets, while enabling interpretable reconstruction of manifold dimensionality, curvature, and coordinate systems.
Deep generative models pose privacy and compliance risks due to unintended memorization of training data. To address this, we propose the Manifold Memorization Hypothesis (MMH), establishing the first geometric unifying framework for memorization—grounded in the dimensional relationship between data manifolds and learned model manifolds. MMH formally defines memorization strength and rigorously distinguishes two distinct mechanisms: overfitting-driven memorization and distribution-driven memorization. Through manifold dimension estimation, synthetic data modeling, and systematic empirical evaluation on large-scale image models—including Stable Diffusion—we validate MMH’s explanatory power for observed memorization phenomena. Furthermore, we develop scalable memorization detection and suppression methods, demonstrating their effectiveness on both synthetic and real-world image datasets. This work provides both a theoretical foundation and practical tools for enhancing privacy safety in generative AI systems.
Existing methods for generating manifold meshes typically rely on indirect representations—such as level sets or template deformations—making it difficult to directly produce high-quality, topologically unconstrained polygonal meshes with structural integrity. This paper introduces the first end-to-end differentiable framework that explicitly models half-edge structure via vertex-level continuous connectivity embeddings, enabling direct generation of discrete manifold-conforming meshes in a continuous latent space. Key contributions include: (1) the first continuous neighborhood relation learning mechanism; (2) mesh distribution fitting via stochastic optimization; and (3) topology-agnostic generation and repair capabilities. Evaluated on large-scale datasets, our method significantly improves mesh element quality, geometric fidelity, and topological diversity. It establishes the first truly end-to-end differentiable approach for manifold mesh generation and repair, bridging a critical gap between implicit representation learning and explicit, valid mesh synthesis.
Traditional machine learning struggles to effectively model shape data with nonlinear geometric structures and their intrinsic variability. This work proposes a unified analytical framework that systematically integrates differential geometry, manifold statistics, and geometric deep learning to address the challenges posed by complex, unaligned shapes exhibiting nonlinear variation. The framework encompasses key components including shape representation, geodesic metrics, parametrization, and statistical inference. It has been successfully applied to multiscale biological geometric data—such as cellular morphologies and primate dental evolution—revealing structural patterns and evolutionary trajectories underlying shape variation. This approach establishes both a theoretical foundation and a practical paradigm for geometry-aware learning in shape analysis.
This work addresses the lack of a unified theoretical framework for assessing the reliability of nonlinear dimensionality reduction embeddings. It proposes a cohesive perspective grounded in differential and integral geometry, systematically analyzing the geometric properties of differentiable embeddings through both local differential structure and global path integrals. The study reveals, for the first time, that multiple existing diagnostic methods fundamentally arise from a common geometric object and demonstrates that their global characteristics cannot be fully captured by derivatives of any finite order, necessitating an irreducible integral viewpoint. Leveraging tools such as curvature analysis, path-dependence detection, and mapping continuity evaluation, the framework validates its theoretical predictions on both synthetic and real-world datasets—including single-cell data—enabling precise estimation of embedding reliability and effective discrimination between single-valued and path-dependent embeddings.
This work addresses the semantic discontinuities that arise during editing in latent diffusion models, which stem from the structural fragility of the latent space. The authors propose a Riemannian geometric framework that decouples the generative Jacobian into local scaling (capacity) and local curvature (complexity). Through this decomposition, they reveal that out-of-distribution generation erroneously allocates curvature to unstable semantic boundaries rather than perceptual details. The study introduces “geometric hotspots” as an intrinsic diagnostic metric to pinpoint the structural origins of generation instability. This approach provides a robust, geometry-aware measure for evaluating and enhancing the reliability of generative models.
This work addresses the challenges of modeling complex dependencies and mitigating redundancy in high-dimensional parameter spaces for discrete data generation. It introduces, for the first time, a Riemannian geometric structure with isometric properties into the exponential parameter space of product manifolds over categorical distributions, thereby constructing a low-dimensional latent subspace. By leveraging the Riemannian metric, geodesics within this subspace become straight lines, enabling consistent and efficient flow-matching training. The proposed approach substantially reduces the dimensionality of latent variables while preserving strong representational capacity for discrete data distributions. Experimental results demonstrate that the model achieves accurate and efficient discrete data generation using a significantly lower-dimensional latent space, effectively balancing computational efficiency with modeling performance.
This work investigates the local geometric structure of controlled semantically similar samples in sentence embedding spaces. It proposes a geometric-aware representation analysis framework that fits low-degree polynomial surfaces—affine, quadratic, and cubic—to local PCA subspaces, incorporating Hessian-based shape descriptors and synthetic point generation in the latent space. The study introduces CoPaGE-300K, a large-scale dataset of controllable variants with slot annotations. Experimental results demonstrate that nonlinear local models more accurately capture the underlying local manifold than affine approximations. While synthetically generated points exhibit strong geometric consistency with the local structure, they do not directly improve downstream classification performance, revealing a fundamental distinction between geometric fidelity and discriminative utility in representation learning.