Score
Designs and analyzes geometric structures on spaces of probability distributions and model families—constructing nonparametric statistical manifolds and Orlicz-function–based fiber-bundle representations that encode total system state and separate structural (model) from statistical (noise/variance) directions. Equips these manifolds with information‑geometric metrics (e.g., the Fisher information metric), computes distances and curvature, and uses those quantities to standardize comparisons across populations, relate geometric distance to likelihood comparisons, and predict curvature-driven dynamics such as boundary or focal formation that inform generalization and model behavior.
In infinite-dimensional nonparametric information geometry, the Fisher–Rao metric’s functional nature renders its inversion intractable—long posing a “computational intractability barrier.” Method: We propose an orthogonal decomposition framework on the tangent space, yielding a finite-dimensional, computable covariate Fisher information matrix (cFIM) under observable covariates. This integrates covariate projection with second-order curvature analysis of the KL divergence within a semiparametric modeling paradigm. Contributions: (i) We prove that the G-entropy equals the trace of cFIM, establishing it as a geometric invariant; (ii) we strengthen the manifold assumption into a testable condition—rank deficiency of cFIM; (iii) we define the information capture ratio, enabling rigorous intrinsic dimension estimation. Our framework provides geometrically grounded, statistically valid tools for quantifying both statistical coverage and model efficiency in explainable AI, and enables verifiable intrinsic dimension inference in high-dimensional settings.
This work addresses the limitations of classical Fisher–Rao information geometry in handling non-regular statistical models—such as those lacking densities, featuring parameter-dependent support, or exhibiting biased score functions—by developing a unified weak information geometry framework. Within the setting of tempered distributions, the authors introduce weakly regular inference functions and Stein representations as tools for information extraction, constructing a Riemannian metric induced by Godambe information that transcends Fisher–Rao constraints. This approach yields a coherent geometric theory for five classes of non-regular settings, including likelihood-free models, varying-support families, transformation-based inference, and mixture models, while revealing a geometric correspondence between inferential non-identifiability and block-diagonal metric structure. Leveraging tempered distribution theory, Schwartz kernels, quadratic Stein discrepancies, and reproducing kernel Hilbert spaces, the study derives closed-form weak Godambe metrics for models such as Cantor location families, uniform scale families, shifted exponentials, hierarchical mixtures, and α-stable noise-driven lattice heat equations, demonstrating their geometric validity and stability.
This work addresses the limitation of classical Fisher information asymptotics, which captures only the first-order geometry of parameter estimation covariance and fails to accurately characterize bias in finite samples. By viewing regular parametric families as Riemannian manifolds equipped with the Fisher–Rao metric and embedding square-root densities into an L² space, the authors derive a second-order (n⁻²) correction term for the covariance. They innovatively unify intrinsic Ricci curvature, extrinsic second fundamental form, and the Hellinger divergence tensor to construct a coordinate-invariant, curvature-aware framework for higher-order covariance expansion, extending it to singular statistical models. This geometric approach elucidates the mechanisms linking learning rates and posterior mean squared error, offering new principles for diagnosing and optimizing weak identifiability.
This work addresses the limitation of existing high-dimensional complexity measures—such as Gaussian width—which rely on Euclidean geometry and fail to capture the intrinsic Riemannian structure and parametrization invariance of statistical models. We propose Fisher width, a generalization of Gaussian width to the statistical manifold endowed with the Fisher information metric, by rescaling sets via the local metric tensor to reflect statistical distinguishability. This measure provides the first parametrization-invariant geometric complexity tool for statistical manifolds, capturing anisotropic geometric effects overlooked by Euclidean approaches. Leveraging differential geometry, information geometry, and high-dimensional probability, we establish concentration properties, perturbation stability, and spectral comparison bounds for Fisher width, and design a computable estimator. Under a Fisher-Lipschitz assumption, we derive generalization bounds and empirically validate the method’s efficacy across three model classes on MNIST.
Existing discrete data generation methods lack efficient and accurate modeling of categorical distributions. Method: This paper proposes the first flow-matching framework grounded in information geometry. It constructs a Riemannian structure on the categorical statistical manifold via the Fisher–Rao metric and defines optimal transport paths along geodesics, enabling exact likelihood computation without variational lower-bound constraints. Crucially, it integrates information geometry with flow matching for the first time, eliminating reliance on simplistic priors (e.g., uniform or independent distributions). Contribution/Results: We develop a natural-gradient-driven geodesic flow-matching algorithm that supports diffeomorphic mappings and invertible modeling. Empirically, our method substantially outperforms state-of-the-art discrete diffusion and flow models on image, text, and biological sequence generation tasks—simultaneously improving both sample quality and log-likelihood accuracy.
Traditional machine learning struggles to effectively model shape data with nonlinear geometric structures and their intrinsic variability. This work proposes a unified analytical framework that systematically integrates differential geometry, manifold statistics, and geometric deep learning to address the challenges posed by complex, unaligned shapes exhibiting nonlinear variation. The framework encompasses key components including shape representation, geodesic metrics, parametrization, and statistical inference. It has been successfully applied to multiscale biological geometric data—such as cellular morphologies and primate dental evolution—revealing structural patterns and evolutionary trajectories underlying shape variation. This approach establishes both a theoretical foundation and a practical paradigm for geometry-aware learning in shape analysis.
This work proposes a unified, imaging-modality-agnostic geometric framework to characterize the information structure of imaging operators. By mapping the normalized singular spectrum onto a probability simplex and introducing the Fisher–Rao information metric, the authors construct a Riemannian geometry of constant positive curvature. This approach yields, for the first time, an intrinsic description at the operator level that does not rely on optimization, stochastic modeling, or modality-specific assumptions, thereby revealing the geometric invariance of spectral equivalence classes. The study derives closed-form expressions for information distance and geodesics, proves their invariance under unitary transformations and global scaling, and elucidates the mechanism by which nonlinear redistribution of spectral weights influences information distance.
This work addresses the lack of effective support for Kendall’s 3D shape space in existing Python libraries—such as Geomstats—which has hindered the practical application of Riemannian geometry in three-dimensional shape analysis. We present the first systematic implementation of an efficient and user-friendly Python toolkit tailored specifically to Kendall’s 3D shape space, enabling scale-, position-, and orientation-invariant shape modeling and statistical analysis. By providing a ready-to-use, open-source solution for shape statistics on manifolds, this contribution fills a critical software gap in advanced 3D shape analysis, substantially lowering the barrier to entry for researchers, improving computational efficiency, and enhancing the reproducibility of results.
This work proposes a Statistical Meaning Geometry (SMG) framework to rigorously distinguish whether large models exhibit genuine intelligence or merely rely on statistical pattern matching. By modeling over-parameterized learning systems as infinite-dimensional Orlicz fiber bundles, the framework introduces a nonlinear curvature-driven mechanism of gauge symmetry breaking, integrated with structural G-entropy, a minimal energy path criterion, and causal invariance filtering to formulate a parameter-free, falsifiable criterion for the emergence of intelligence. Under out-of-distribution (OOD) stimuli, experiments reveal an integer-order +1.0 jump in G-entropy—a signature that enables, for the first time, mathematically rigorous certification of autonomous scientific discovery and paradigm shifts.
Traditional linear dimensionality reduction methods often fail to effectively uncover the intrinsic low-dimensional manifold structure embedded in high-dimensional data. This work systematically traces the historical development of manifold fitting and, for the first time, categorizes it into three distinct phases: nonparametric statistics, mathematically inspired analysis, and modern practical statistics. It clarifies manifold fitting’s role as an independent geometric data analysis tool and delineates its conceptual boundaries from related techniques such as manifold embedding and denoising. By integrating nonparametric methods, differential geometry, and contemporary statistical learning approaches, the paper explores cutting-edge applications of manifold fitting in neural networks and bioinformatics, offering a comprehensive reference framework that elucidates both its theoretical limits and practical utility.