🤖 AI Summary
Inference on within- and between-group effects in high-dimensional repeated-measures data (where dimension $d geq N$) remains challenging due to reliance on restrictive assumptions about covariance structure and the $d/N$ asymptotic regime.
Method: We propose a parametric inference framework that avoids such assumptions, constructing exact test statistics under a multivariate normal model. The approach integrates an efficient shrinkage-based covariance estimator with a randomized subsampling acceleration strategy, substantially reducing computational complexity.
Contribution/Results: We develop hdrm—an open-source R package—capable of heterogeneous-covariance multi-group comparisons for the first time. The framework unifies single- and multi-group settings, as well as homoscedastic and heteroscedastic covariance scenarios, while preserving statistical validity. It extends the applicability boundary of high-dimensional longitudinal data analysis and establishes a new paradigm for robust inference in biomedical and psychological research.
📝 Abstract
Repeated-measure designs allow comparisons within a group as well as between groups, and are commonly referred to as split-plot designs. While originating in agricultural experiments, they are now widely used in medical research, psychology, and the life sciences, where repeated observations on the same subject are essential.
Modern data collection often produces observation vectors with dimension $d$ comparable to or exceeding the sample size $N$. Although this can be advantageous in terms of cost efficiency, ethical considerations, and the study of rare diseases, it poses substantial challenges for statistical inference.
Parametric methods based on multivariate normality provide a flexible framework that avoids restrictive assumptions on covariance structures or on the asymptotic relationship between $d$ and $N$. Within this framework, the freely available R-package hdrm enables the analysis of a wide range of hypotheses concerning expectation vectors in high-dimensional repeated-measure designs, covering both single-group and multi-group settings with homogeneous or heterogeneous covariance matrices.
This paper describes the implemented tests, demonstrates their use through examples, and discusses their applicability in practical high-dimensional data scenarios. To address computational challenges arising for large $d$, the package incorporates efficient estimators and subsampling strategies that substantially reduce computation time while preserving statistical validity.