Score
Designs and implements cohort definitions and data pipelines to select, construct, curate, harmonize, and validate groups of subjects across datasets and time, including specification of inclusion/exclusion criteria, entry and exit windows, and stability checks. Builds and applies longitudinal and cross‑cohort analyses — e.g., survival and entry‑cohort analysis, cohort segmentation, age‑period‑cohort (APC) decomposition, selection‑versus‑adaptation tests, and cross‑cohort validation — to separate age/period/cohort influences, isolate lifecycle or merger effects, estimate time‑varying responses, and compare predictive performance across cohorts.
Longitudinal omics data pose significant challenges for dynamic modeling and clinical translation due to their high dimensionality, temporal imbalance, and non-Gaussian distributional properties. To address these challenges, this study rigorously delineates the theoretical boundaries between time-series and longitudinal analysis, and establishes a methodology classification framework specifically tailored to omics characteristics—encompassing single-cell longitudinal modeling, multi-omics integration, network dynamics, and FDA-compliant analysis. We systematically unify linear and generalized linear mixed models, functional data analysis, Bayesian hierarchical modeling, survival analysis, and multi-view cross-platform fusion algorithms. The resulting methodological guide comprehensively addresses modeling assumptions, algorithmic suitability, and software implementation, delivering a reproducible, scalable analytical framework. This work substantially enhances the rigor, interpretability, and translational utility of complex longitudinal omics studies.
Existing dynamic survival prediction methods lack standardized benchmarking, particularly for high-dimensional longitudinal biomedical data. Method: This study systematically evaluates multi-step dynamic survival prediction approaches—including mixed-effects models, multivariate functional principal component analysis, Cox regression, random survival forests, and landmark analysis—across multiple real-world datasets. We rigorously control key factors (sample size, covariate dimensionality, and follow-up duration) to assess predictive accuracy, robustness, and computational efficiency. Contribution/Results: The analysis reveals critical trade-offs among modeling choices, identifies context-specific adaptability requirements, and delineates practical performance limits. It provides the first empirically grounded, methodological guideline for selecting appropriate dynamic risk prediction strategies in clinical settings, thereby advancing evidence-based decision support for time-varying prognostic modeling.
Medical machine learning faces irreproducibility in task definition and cohort construction due to the privatization of electronic health record (EHR) data. Method: We propose an automated cohort extraction system for event-stream EHR data, introducing the first domain-specific configuration language (DSL) designed specifically for event streams. This DSL decouples general inclusion/exclusion criteria from dataset-specific clinical concepts. The system supports zero-code adaptation to standardized EHR formats (e.g., MEDS, ESGPT) via a modular pipeline integrating DSL compilation, event-stream parsing, rule-driven cohort generation, and standardized interface adapters. Contribution/Results: Experiments across multi-institutional real-world EHR datasets demonstrate consistent cohort definitions across formats and institutions. Our approach significantly lowers the barrier to task specification, enables concept-level reproducibility across datasets, ensures precise cohort replication within a dataset, and enhances reproducibility and collaborative efficiency in EHR-based research.
Longitudinal data often exhibit multiple sources of heterogeneity, including divergent mean trajectories, increasing residual variance over time, and occasional outlying measurements. Conventional homogeneous models may yield inefficient parameter estimates and inflated variance assessments in such settings. This work proposes a novel Bayesian mixture model that, for the first time, incorporates covariate-driven binary indicator variables within a unified Bayesian framework to jointly model these three forms of heterogeneity via logistic regression. Inference is carried out using Markov chain Monte Carlo (MCMC) methods, and the approach facilitates posterior-probability-based model selection to evaluate the necessity of each heterogeneous component. Simulation studies demonstrate that the proposed method accurately identifies underlying heterogeneity structures and yields efficient fixed-effect estimates. Its practical utility is further corroborated through application to DHEAS hormone data from the Study of Women’s Health Across the Nation (SWAN).
This study addresses the heterogeneity in patients’ clinical trajectories—particularly in diagnostic and therapeutic sequences and their temporal documentation—by proposing a two-stage interpretable modeling framework. First, a rule-based algorithm extracts representative care pathways from cohort data; these pathways are then formalized as Markov chains and integrated as mixture components into a weighted probabilistic model of individual patient trajectories. This approach innovatively decouples population-level pathway discovery from patient-specific modeling, enabling each patient’s trajectory to be represented as a probabilistic combination of multiple canonical pathways. The method enhances both interpretability and clinical utility, as demonstrated on real-world data from radical prostatectomy patients, where it consistently identified clinically meaningful care patterns and effectively facilitated patient subgroup discovery.
This study addresses a key limitation of conventional joint models, which focus solely on the mean trajectories of biomarkers while ignoring the prognostic value embedded in within-individual variability. To overcome this, the authors propose a two-step approach: first, individual- and time-specific variability metrics are derived from residuals of a mixed-effects model; second, these metrics are incorporated into a standard joint modeling framework to simultaneously assess the effects of both mean levels and variability on survival outcomes. The method requires no specialized software and can be flexibly integrated into existing joint modeling platforms such as JM or joineR, with support for multiple biomarkers. Simulation studies demonstrate robust performance across diverse scenarios, and application to glioblastoma clinical data reveals that both the mean and variability of white blood cell counts are significantly associated with overall survival.
This study addresses the efficient integration of individual participant data (IPD) and aggregate data (AD) under dataset shift, systematically evaluating the impact of different AD formats on estimation efficiency within a constrained maximum likelihood framework. The authors propose a fast, non-iterative estimation algorithm robust to both covariate shift and prior probability shift. Their analysis reveals that outcome-stratified summary statistics—such as case/control counts—substantially outperform covariate-stratified summaries, particularly in settings with continuous outcomes, yielding markedly improved estimation efficiency. The method’s stability, scalability, and computational advantages are demonstrated through empirical validation on income data (exhibiting covariate shift) and housing data (exhibiting prior probability shift), offering theoretical support for standardized reporting of AD in clinical trial publications.