Score
Designs and implements methods to represent and summarize functional data by extracting functional principal components (including covariate-adjusted FPCA) and producing low-dimensional scores that capture dominant modes of inter-individual variation. Builds regression models that relate those component scores or the observed functions to scalar or functional outcomes, adjusting for covariates and enabling adjusted or counterfactual functional estimates.
Non-random missingness due to detection limits—common in longitudinal biomarkers and other functional data—constitutes a missing-not-at-random (MNAR) mechanism. Conventional functional principal component analysis (FPCA) imputes values at detection limits, ignoring the MNAR structure and inducing bias in eigenfunction and score estimation. Method: We propose the first FPCA framework specifically designed for MNAR functional data, systematically integrating detection-limit–adjusted mean and covariance function estimation directly into the FPCA pipeline—thereby circumventing imputation-induced bias. Contribution/Results: Building on the estimation theory of Liu & Houwing-Duistermaat (2022, 2023), our method achieves consistency and robustness under both sparse and dense sampling regimes, as confirmed by asymptotic analysis and extensive simulations. Empirical evaluation—including simulation studies and real-data applications—demonstrates substantially improved accuracy in eigenfunction and score estimation, enhancing the reliability and applicability of functional data analysis under detection limits.
This study addresses the challenge of extracting principal components from sparse and irregularly observed multivariate functional data, particularly in longitudinal settings where modeling cross-variable dependencies is difficult. To overcome the limitations of conventional approaches that rely on univariate scores and subsequent eigendecomposition, the authors propose a novel framework that directly estimates multivariate functional principal components by integrating maximum likelihood estimation with a modified Gram–Schmidt orthogonality constraint. This approach more accurately captures the covariance structure among variables. Empirical evaluations on datasets comprising Alzheimer’s disease cognitive biomarkers and Irish dairy cow milk production demonstrate that the proposed method yields substantially improved estimation accuracy and interpretability of principal components compared to existing techniques.
This study addresses the theoretical gap in understanding why functional principal component analysis (FPCA) can severely fail when applied to rough functional data. The authors propose a model that explicitly characterizes data roughness and, for the first time, theoretically elucidate the mechanism through which roughness induces bias in FPCA. They further identify a phase-transition threshold governing the loss of information in FPCA. By integrating tools from random matrix theory, generic chaining techniques, and functional data analysis, they derive spectral statistics suitable for model diagnostics and goodness-of-fit testing. Comprehensive theoretical analysis, simulations, and empirical applications to climate and environmental data demonstrate that the proposed method accurately delineates the performance boundaries of FPCA and provides a practical diagnostic tool for assessing the validity of estimated principal components.
This study addresses the challenges of functional principal component analysis (FPCA) for sparse functional data—characterized by few irregularly spaced observations and measurement error—by proposing a novel approach that integrates basis expansion with maximum likelihood estimation. The method explicitly enforces orthogonality of eigenfunctions during optimization via a modified Gram–Schmidt procedure and reconstructs complete trajectories using conditional expectation estimates of principal component scores. A key innovation is the introduction of an information criterion that simultaneously selects the number of basis functions and the rank of the covariance operator, enabling adaptive model selection. Extensive simulations and real-data analyses on CD4 cell counts and dairy cow somatic cell counts demonstrate that the proposed method substantially outperforms existing approaches in estimating principal components and reconstructing trajectories under sparsity.
Traditional functional principal component analysis (FPCA) treats principal component estimates as deterministic quantities, neglecting their sampling variability and thereby yielding inaccurate uncertainty quantification. To address this, we propose Bayes-FPCA—a fast Bayesian FPCA framework that models principal components directly on the Stiefel manifold for the first time. It achieves efficient dimensionality reduction via orthogonal spline basis projection and constructs a uniform prior on the manifold using the polar decomposition. A stable MCMC sampling strategy is designed to respect the structural constraints inherent in FPCA. Evaluated on DASH4D continuous glucose monitoring data, Bayes-FPCA accurately characterizes postprandial glucose dynamics and substantially improves uncertainty estimation accuracy while accelerating computation by an order of magnitude. All code and simulation routines are publicly available.
This study addresses a key limitation of existing multivariate functional principal component analysis (MFPCA) methods, which assume all functional observations share a common domain and thus struggle with real-world data exhibiting varying domains. To overcome this, the authors propose the first MFPCA framework tailored for variable-domain settings. Their approach first performs univariate FPCA separately for each variable over its individual domain, stacks the resulting scores, and then smooths the empirical covariance matrix according to domain lengths to estimate multivariate eigenfunctions and scores that properly account for heterogeneous observation intervals. By explicitly incorporating individual domain information, the method avoids biases introduced by conventional binning or by ignoring domain variability. Simulation studies demonstrate its superior performance over existing strategies, and the approach is successfully applied to analyze temperature and blood oxygen saturation trajectories in COVID-19 patients.
This study addresses dimensionality reduction and covariance analysis for scalar-valued functional data indexed over a compact d-dimensional Riemannian manifold, where the indexing domain—not the function values—exhibits geometric structure. The work extends functional principal component analysis (FPCA) to manifold-indexed settings by introducing an intrinsic kernel estimation framework that incorporates geodesic distance and Riemannian volume density correction, accommodating heterogeneous sampling frequencies and weighted estimation strategies across individuals. Theoretically, by integrating intrinsic kernel methods, VC-type empirical process theory, and spectral perturbation analysis, the paper establishes uniform convergence rates for the mean, covariance, and eigen-objects under non-Lipschitz kernels. It further reveals that the transition from sparse to dense observation regimes is governed by the intrinsic manifold dimension, reducing to classical results when d = 1. Experiments on S¹, S², and real-world SONICOM head-related transfer function data demonstrate consistent and substantial performance gains over baseline methods that ignore geometric structure.
Traditional approaches often reduce physical activity data to scalar summaries, failing to capture the time-varying dynamics of intervention effects. This study addresses this limitation by employing a functional regression framework that treats entire activity trajectories as functional observations, thereby directly modeling the temporal evolution of intervention effects. By replacing conventional two-stage methods with function-on-scalar regression (FoSR) and further extending it to function-on-function regression (FoFR), combined with functional principal component analysis (FPCA), the approach substantially enhances the ability to analyze high-dimensional outcome data. Applied to the STEP UP trial, this method successfully identified three distinct intervention strategies that exerted significant, interpretable, and differentially sustained time-varying effects on daily step counts.
This study addresses critical limitations of regression on principal component scores (RPCS) in associating functional data with scalar covariates, including loss of statistical power, poor control of Type I error, and lack of valid inference. Through a systematic comparison with function-on-scalar regression (FoSR), the work quantifies for the first time the mechanism underlying RPCS’s power loss, demonstrating its dependence on the correlation between principal components and the true effect function. The authors further establish that existing RPCS approaches generally fail to yield reliable statistical inference. Extensive simulations and an analysis of minute-level accelerometer data from NHANES confirm that FoSR consistently outperforms RPCS in terms of power, Type I error control, and estimation accuracy, thereby establishing FoSR as a more robust and trustworthy alternative for functional data analysis.