Score
Designs and implements functional principal component analysis methods for data consisting of functions whose domain is a manifold, including construction and spectral decomposition of the manifold-indexed covariance operator to compute eigenvalues and eigenfunctions and to produce low-dimensional representations (principal component scores and modes) of manifold-indexed functions. Analyzes statistical properties and sampling behavior of these methods, deriving intrinsic, manifold-aware error bounds (e.g., Hilbert-Schmidt and operator-norm rates) and characterizing sparse-to-dense sampling transitions in terms of the manifold's intrinsic dimension.
This study addresses dimensionality reduction and covariance analysis for scalar-valued functional data indexed over a compact d-dimensional Riemannian manifold, where the indexing domain—not the function values—exhibits geometric structure. The work extends functional principal component analysis (FPCA) to manifold-indexed settings by introducing an intrinsic kernel estimation framework that incorporates geodesic distance and Riemannian volume density correction, accommodating heterogeneous sampling frequencies and weighted estimation strategies across individuals. Theoretically, by integrating intrinsic kernel methods, VC-type empirical process theory, and spectral perturbation analysis, the paper establishes uniform convergence rates for the mean, covariance, and eigen-objects under non-Lipschitz kernels. It further reveals that the transition from sparse to dense observation regimes is governed by the intrinsic manifold dimension, reducing to classical results when d = 1. Experiments on S¹, S², and real-world SONICOM head-related transfer function data demonstrate consistent and substantial performance gains over baseline methods that ignore geometric structure.
This paper addresses the challenging problem of robustly estimating the true rank—i.e., the number of principal components—of the covariance operator for discretely observed, noisy functional data. Measurement error induces spurious ridge effects in the empirical covariance function, severely complicating rank determination. We propose a novel method that jointly leverages eigenvalue decay rates and eigenfunction smoothness, integrating smoothed nonparametric covariance estimation, adaptive thresholding, and simulation-based calibration. The approach accommodates random and subject-specific sampling designs. Theoretically, it guarantees strong robustness against noise; simulations demonstrate substantially higher accuracy than classical information criteria (e.g., AIC, BIC), maintaining stability even under high noise levels. Empirical validation on two real functional datasets confirms its practical utility and effectiveness.
Traditional functional principal component analysis (FPCA) yields globally supported eigenfunctions, resulting in poor interpretability. To address this, we propose localized functional principal component analysis (LFPCA), a covariance-structure-driven piecewise decomposition method that orthogonally decomposes a stochastic process into statistically independent sub-processes with disjoint supports. FPCA is then applied independently on each subinterval, yielding orthogonal eigenfunctions with compact (local) support. Crucially, LFPCA achieves localization without sparse regularization, rigorously preserving both the original covariance structure and orthogonality of the decomposition. Moreover, the variance contribution of each sub-process is directly quantified by its associated eigenvalue, eliminating structural distortion caused by over-regularization. Extensive experiments on synthetic and real-world data demonstrate that LFPCA significantly improves interpretability, fidelity to orthogonality, and accuracy in variance explanation compared to conventional FPCA.
This paper addresses the problem of functional principal component estimation from noisy, discretely sampled functional data, with the goal of characterizing the statistical impact of denoising preprocessing. Under a double-asymptotic regime where both sample size and number of observation grid points grow, we propose a histogram-projection-based estimator for functional principal components. We establish, for the first time under realistic joint constraints of noise and discrete sampling, the minimax optimal convergence rate for functional principal component estimation and rigorously prove that our method achieves this rate. Theoretically, we show that smoothing preprocessing improves the convergence order but does not alter the fundamental minimax difficulty. Extensive simulations validate the method’s effectiveness, and we demonstrate its practical utility through functional visualization analysis of genomic data.
In principal component analysis (PCA), near-degenerate eigenvalues induce the “isotropy curse,” causing instability in principal direction estimation and diminished interpretability. Method: This paper proposes a novel modeling framework leveraging the eigenvalue multiplicity hierarchy of the covariance matrix. Generalizing probabilistic PCA (PPCA), it employs the ordered geometric structure of flag manifolds to characterize maximum-likelihood estimation under joint eigenvalue multiplicity constraints on signal and noise subspaces. Contribution/Results: We introduce, for the first time, a hierarchical partial-order model selection criterion enabling compact, interpretable low-dimensional modeling. Experiments demonstrate that our method significantly outperforms PPCA—particularly in small-sample regimes and when eigenvalue gaps are weak—achieving superior trade-offs between model complexity and fitting accuracy on both synthetic and real-world data.
This study addresses the theoretical gap in understanding why functional principal component analysis (FPCA) can severely fail when applied to rough functional data. The authors propose a model that explicitly characterizes data roughness and, for the first time, theoretically elucidate the mechanism through which roughness induces bias in FPCA. They further identify a phase-transition threshold governing the loss of information in FPCA. By integrating tools from random matrix theory, generic chaining techniques, and functional data analysis, they derive spectral statistics suitable for model diagnostics and goodness-of-fit testing. Comprehensive theoretical analysis, simulations, and empirical applications to climate and environmental data demonstrate that the proposed method accurately delineates the performance boundaries of FPCA and provides a practical diagnostic tool for assessing the validity of estimated principal components.
This study addresses the well-posedness of the posterior distribution in fully Bayesian functional principal component analysis (FPCA). By projecting functions onto a spline basis, functional orthogonality is recast as an orthogonality constraint on the coefficients, and a smoothing penalty based on the integral of the squared second derivative is incorporated. The smoothing parameter is treated as the inverse variance component in a mixed-effects model. The work innovatively formulates the smoothing penalty through the eigenvalues of its design matrix and, for the first time, establishes a sufficient condition for posterior well-posedness. This result yields a simple and practical criterion for selecting priors on the smoothing parameter, thereby substantially enhancing the stability and interpretability of Bayesian FPCA models.
This study investigates how to effectively transfer the supremum-norm convergence rates and weak convergence properties of covariance kernel estimators to functional principal components (FPCs). By integrating $L_2$ perturbation theory with the Nyström method and minimax lower bound analysis, the work establishes, for the first time, optimal supremum-norm convergence rates and asymptotic normality for FPC estimators. The analysis uncovers a novel phenomenon in sparse observation regimes, where eigenvalue estimation is dominated by discretization effects. These theoretical results are further extended to cross-covariance, long-run covariance, and derivative process covariance kernels. Empirical validation on temperature curves demonstrates the method’s effectiveness across the spectrum from sparse to dense observation designs.