Score
Applying within-subjects statistical analysis to detect moderators and quantify temporal and between-seed variability, backslides, spikes, and rolling volatility across intervals when outcomes are measured repeatedly.
Existing methods for causal mediation analysis face limitations in handling continuous or multidimensional mediators, non-binary treatments, and intermediate confounding, and often lack usability. This work proposes the crumble framework, which, for the first time, provides a unified nonparametric approach to estimating diverse mediation effects—including natural direct and indirect effects and stochastic intervention effects—under intermediate confounding, while accommodating treatment variables of arbitrary type. Built upon semiparametric theory and leveraging modified treatment policies, crumble enables flexible modeling with high interpretability. The framework’s robustness, practicality, and broad applicability are demonstrated through two real-data applications involving both binary and non-binary treatments.
Existing graph-based change-point detection methods perform well in high-dimensional nonparametric settings but struggle with data featuring repeated measurements or local group structures; conventional mean aggregation obscures intra-individual dynamic patterns. This paper proposes the first graph-based framework that jointly leverages intra- and inter-individual information: it constructs a two-layer similarity graph—comprising an intra-individual graph capturing repeated measurements and an inter-individual proximity graph—thereby integrating local dependencies with global structure. We design a composite test statistic and derive its p-value efficiently via analytic approximation. The method achieves high sensitivity to subtle, gradual, and multi-scale change points. Experiments on New York City taxi trajectory data demonstrate substantial improvements in detection accuracy and robustness over state-of-the-art approaches, particularly under strong inter-individual heterogeneity.
This study addresses the unclear mechanisms through which time-varying covariates (TVCs) influence latent class trajectory heterogeneity in nonlinear growth mixture models (GMMs). We propose a novel TVC decoupling framework that decomposes each TVC into two distinct components: a baseline trait—capturing its effect on initial growth factors—and a time-specific state—directly influencing observed outcomes—while allowing heterogeneous effects across latent classes. Methodologically, we integrate GMM with an extended mixture-of-experts (MoE) architecture and the proposed TVC decomposition, validated via Monte Carlo simulations and empirical longitudinal analysis. Simulation results confirm unbiased parameter estimation and nominal coverage of confidence intervals. Empirically, we identify significant between-class heterogeneity in both baseline and dynamic effects of reading ability on mathematics achievement trajectories. To our knowledge, this is the first work to systematically disentangle baseline versus dynamic TVC effects and model their cross-class heterogeneity within nonlinear GMMs, thereby enhancing theoretical precision and empirical interpretability in attributing trajectory heterogeneity.
Existing structural equation modeling (SEM) frameworks struggle to model latent variable variances that depend on other latent variables, thereby limiting the characterization of latent heteroscedasticity—such as in psychological constructs like personality or creativity. To address this, we propose Bayesian Gaussian Distributional SEM, the first SEM extension integrating distributional regression into the SEM framework to jointly model both the mean and variance of latent variables. Leveraging Bayesian inference and MCMC sampling, our approach flexibly specifies latent variances as arbitrary functions of other latent variables. Simulation studies demonstrate high statistical reliability and computational efficiency. Empirical analysis of personality data reveals that emotional stability significantly moderates the variability of neuroticism—a finding inaccessible under conventional SEM. This work introduces a novel theoretical tool and methodological paradigm for modeling latent heteroscedasticity, advancing both substantive theory testing and statistical methodology in behavioral and social sciences.
This study addresses the challenge of modeling multivariate longitudinal data characterized by individually varying measurement occasions, nonlinear developmental trajectories, and interdependent outcomes. We propose the Multivariate Latent Baseline Growth Model (MLBGM), a structural equation modeling–based framework that extends the latent baseline growth concept to joint modeling of multiple outcomes. MLBGM explicitly incorporates person-specific measurement times, thereby relaxing the restrictive assumption of equidistant assessments inherent in conventional growth models. Monte Carlo simulations demonstrate unbiased parameter estimates and adequate confidence interval coverage. Applied to large-scale educational longitudinal data, MLBGM uncovers asynchronous, nonlinear, and coupled developmental patterns between literacy and mathematics competencies. Our key contributions are: (1) the first latent growth model accommodating both individually varying measurement times and multidimensional nonlinear co-evolution; and (2) publicly available open-source implementation code.
Existing online learning analytics struggle to capture the nonlinear dynamics of student ability, often requiring a prespecified number of clusters and failing to effectively model the relationship between engagement behaviors and ability evolution. This work proposes a Bayesian nonparametric dynamic item response theory framework that employs B-spline basis functions to flexibly characterize the nonlinear influence of participation on ability drift. By incorporating a Mixture-of-Finite-Mixtures prior, the model automatically infers the number of latent learner subgroups, enabling unsupervised clustering and longitudinal tracking of individual ability trajectories. Applied to data from 198 undergraduate students in a statistics course, the model identified four distinct learner types—struggling-declining, low-stable, mainstream-stable, and high-improving—revealing highly stable ability trajectories and no significant predictive effect of participation volume on ability drift.
This work addresses the challenge of modeling highly heterogeneous disease progression in irregularly sampled, multivariate longitudinal classification data by proposing a continuous latent variable dynamical model that integrates item response theory with a covariate-dependent, non-homogeneous multivariate Ornstein–Uhlenbeck (OU) process. The approach employs a scalable parameterization of the drift matrix to effectively capture coupled evolution across multiple functional domains and supports personalized trajectory modeling with an arbitrary number of latent variables. Theoretical analysis guarantees model identifiability, and experiments on both simulated and real-world ALS clinical data demonstrate the model’s flexibility and effectiveness in characterizing individualized, multi-domain interactive disease dynamics.
This study addresses the challenge of modeling time-varying correlations among multiple longitudinal variables that dynamically depend on individual-level covariates—a feature inadequately captured by existing methods. The authors propose the TiVAC model, which operates within a bivariate Gaussian framework and employs penalized splines in a semiparametric formulation to flexibly characterize the smooth, covariate-dependent evolution of correlations over time. Estimation is efficiently performed via penalized maximum likelihood using a Newton–Raphson algorithm. TiVAC is the first method to enable flexible modeling of covariate-driven time-varying correlation effects while providing simultaneous confidence bands for formal inference. Simulation studies demonstrate its superior performance across diverse scenarios. Application to data from 291 bipolar disorder patients reveals age-dependent heterogeneity in how sex and use of neurologic medications modulate the correlation between depression and anxiety symptoms.
Traditional ecological momentary assessment (EMA) analyses often rely on simplified metrics such as means, which inadequately capture individual longitudinal dynamics. This study moves beyond mean-based approaches by integrating sequence analysis, principal component analysis, and K-means clustering to directly identify latent subgroups exhibiting similar behavioral patterns from full time-series data, while effectively accommodating heterogeneity in sample size and observation frequency. Validated against latent class analysis (LCA) and latent transition analysis (LTA), the proposed method successfully uncovers distinct stress trajectory subgroups in real-world EMA stress data and substantially enhances explanatory power for variability in cognitive performance.
Traditional approaches struggle to capture the temporal evolution of variability, skewness, and tail behavior in the underlying risk distribution of recurrent binary events such as hospital readmissions. This work proposes two novel frameworks: BLaS-Recurrent, a Bayesian model based on the sinh-arcsinh distribution, and QuaD-Recurrent, a quasi-distributional model leveraging nonparametric surface mapping. Both frameworks uniquely embed time-varying location, scale, skewness, and kurtosis into a flexible distributional family, jointly modeling dynamic shifts in both risk level and distributional shape. Moving beyond the limitation of estimating only mean risk, the proposed methods demonstrate superior calibration, robustness, and clinical interpretability in both simulations and real-world MIMIC-IV readmission data, uncovering previously overlooked patterns such as increasing right-skewness and expanding dispersion over time.