Score
Designs, implements, and evaluates statistical and probabilistic methods for datasets with multiple interdependent variables, including estimation of multivariate distributions, computation of multivariate correlations and associations, hypothesis testing, and multivariate regression. Applies dimensionality‑reduction techniques such as principal component analysis and builds multivariate time‑series models to characterize temporal dependencies and predictive relationships among variables.
This study addresses the challenge that conventional dimensionality reduction techniques often fail to preserve critical extremal dependence structures in high-dimensional multivariate extreme value data, thereby compromising the accuracy of subsequent modeling. To overcome this limitation, the authors propose a novel method that integrates principles from extreme value theory with the conceptual framework of principal component analysis, specifically designed to retain extremal dependencies during dimensionality reduction. The proposed algorithm effectively maintains key multivariate extremal characteristics while substantially reducing dimensionality, offering a computationally efficient and statistically accurate approach for analyzing high-dimensional extreme events.
This study addresses the lack of general-purpose modeling and inference frameworks for discrete multivariate time series—such as count, binary, or ordinal categorical data—in fields like psychology and education. The authors propose a copula-based multivariate model grounded in latent-variable Gaussian processes, where discrete observations are generated via deterministic transformations. For the first time, they derive analytic standard errors for the latent Gaussian dynamic parameters within this class of models. Theoretical analysis establishes the asymptotic normality of the joint estimators of the latent processes and marginal distribution parameters. Simulation studies and empirical applications demonstrate that the proposed standard errors perform well in finite samples, enabling reliable statistical inference about the underlying latent Gaussian dynamic structure.
This paper addresses the challenges of covariate effect inference and future outcome prediction in high-dimensional multivariate longitudinal data—characterized by complex dependencies among variables and over time, mixed-type outcomes (continuous and discrete), and irregular observation patterns (missingness or right-censoring). We propose a novel latent-variable modeling framework that (i) unifies the treatment of mixed-type responses and irregular temporal observations for the first time; (ii) introduces an information criterion tailored to high-dimensional longitudinal settings for automatic selection of the latent factor dimension; and (iii) establishes a rigorous central limit theorem for regression coefficient estimators, ensuring valid statistical inference. Evaluated on a customer shopping behavior prediction task, our method significantly improves long-term trend modeling accuracy and robustness of personalized forecasting, demonstrating both practical utility and theoretical soundness in real-world high-dimensional longitudinal applications.
This paper addresses the challenge of jointly modeling covariate effects, spatial dependence, and temporal dependence in multivariate spatiotemporal data. We propose a matrix-response varying-coefficient regression model that embeds covariates into the mean structure of the response matrix and employs a Kronecker-product-based separable covariance structure to explicitly decouple spatial and temporal correlations. Parameter estimation is conducted via maximum likelihood, ensuring both computational efficiency and statistical accuracy, while enabling robust inference under heterogeneous spatial resolutions. Simulation studies demonstrate excellent parameter recovery performance. Applied to municipal-level agricultural and livestock panel data from Brazil, the method successfully uncovers interpretable spatiotemporal dynamic patterns and reveals heterogeneous impacts of key covariates—including climate variables and policy interventions—across space and time. The framework provides a scalable, principled paradigm for high-dimensional spatiotemporal causal analysis.
This paper addresses the challenge of modeling nonlinear and asymmetric dynamic relationships among macroeconomic and financial variables. We propose the first scenario-analysis-oriented, dynamic nonparametric multivariate Bayesian machine learning framework. Methodologically, we adapt classical econometric tools—including conditional forecasting and generalized impulse response analysis—to high-dimensional Bayesian nonparametric models, integrating dynamic factor extensions and Monte Carlo simulation to enable asymmetric shock response estimation and conditional scenario inference. Our key contribution is the first systematic integration of traditional scenario-analysis tools with nonlinear Bayesian machine learning, explicitly capturing structural asymmetry. The framework is validated across three empirical domains: financial stress testing, macroeconomic risk assessment, and cross-border spillover analysis. Results demonstrate substantial improvements in risk measurement accuracy and cross-jurisdictional early-warning capability, offering a novel paradigm for prudential regulation and policy evaluation.
The intrinsic relationship between variable clustering and principal component analysis (PCA) has long been overlooked in the literature. Method: We propose a novel paradigm that applies K-means clustering to the transpose of the data matrix—thereby clustering variables—and quantifies each cluster’s contribution to individual principal components via variable loadings. Contribution/Results: This approach establishes, for the first time, an interpretable mapping between variable clusters and the directions of maximal variance in PCA, yielding a unified “variable clustering–PC contribution” analytical framework. Empirical evaluation demonstrates that the method effectively identifies variable groups driving dominant sources of variation, substantially enhancing interpretability in high-dimensional data. It provides a statistically principled yet computationally feasible tool for multivariate exploratory data analysis.
This study addresses the analysis of multimodal mobile health time-series data comprising continuous, truncated, ordinal, and binary variables by proposing Mixed-type Multivariate Functional Principal Component Analysis (M²FPCA). Built upon a semiparametric Gaussian copula framework, M²FPCA assumes observations arise from an underlying multivariate generalized nonparametric normal functional process. It employs Kendall’s tau bridging to estimate both cross-variable and temporal dependence structures and incorporates a partial separability assumption on the covariance operator to enhance computational efficiency. As the first extension of multivariate functional principal component analysis to mixed data types, M²FPCA yields interpretable latent digital biomarkers. Applied to data from 307 participants, it successfully identifies shared diurnal patterns across mood, anxiety, energy, and physical activity, effectively differentiating subtypes of mood disorders; simulation studies further confirm its superior performance under complex dependency structures.
This study addresses the challenge of extracting principal components from sparse and irregularly observed multivariate functional data, particularly in longitudinal settings where modeling cross-variable dependencies is difficult. To overcome the limitations of conventional approaches that rely on univariate scores and subsequent eigendecomposition, the authors propose a novel framework that directly estimates multivariate functional principal components by integrating maximum likelihood estimation with a modified Gram–Schmidt orthogonality constraint. This approach more accurately captures the covariance structure among variables. Empirical evaluations on datasets comprising Alzheimer’s disease cognitive biomarkers and Irish dairy cow milk production demonstrate that the proposed method yields substantially improved estimation accuracy and interpretability of principal components compared to existing techniques.
We assume that we have multiple ordinal time series and we would like to specify their joint distribution. In general it is difficult to create multivariate distribution that can be easily used to jointly model ordinal variables and the problem becomes even more complex in the case of time series, since we have to take into consideration not only the autocorrelation of each time series and the dependence between time series, but also cross-correlation. Starting from the simplest case of two ordinal time series, we propose using copulas to specify their joint distribution. We extend our approach in higher dimensions, by approximating full likelihood with composite likelihood and especially conditional pairwise likelihood, where each bivariate model is specified by copulas. We suggest maximizing each bivariate model independently to avoid computational issues and synthesize individual estimates using weighted mean. Weights are related to the Hessian matrix of each bivariate model. Simulation studies showed that model fits well under different sample sizes. Forecasting approach is also discussed. A small real data application about unemployment state of different countries of European Union is presented to illustrate our approach.
This study addresses the challenge of modeling time-varying correlations among multiple longitudinal variables that dynamically depend on individual-level covariates—a feature inadequately captured by existing methods. The authors propose the TiVAC model, which operates within a bivariate Gaussian framework and employs penalized splines in a semiparametric formulation to flexibly characterize the smooth, covariate-dependent evolution of correlations over time. Estimation is efficiently performed via penalized maximum likelihood using a Newton–Raphson algorithm. TiVAC is the first method to enable flexible modeling of covariate-driven time-varying correlation effects while providing simultaneous confidence bands for formal inference. Simulation studies demonstrate its superior performance across diverse scenarios. Application to data from 291 bipolar disorder patients reveals age-dependent heterogeneity in how sex and use of neurologic medications modulate the correlation between depression and anxiety symptoms.