🤖 AI Summary
This study addresses a critical gap in genetic research by developing the first closed-form sample size formula for testing the total effect of a single nucleotide polymorphism (SNP)—encompassing both direct effects and indirect effects mediated through longitudinal biomarkers—within joint models of longitudinal and survival data. Existing methods lack appropriate power and sample size calculations for this integrative framework. The proposed approach is grounded in joint modeling theory and rigorously validated via Monte Carlo simulations, demonstrating high accuracy and robustness under finite-sample settings. Simulation studies confirm the reliability of the derived formula, and its practical utility is illustrated through a successful application to data from the Diabetes Control and Complications Trial (DCCT), underscoring its feasibility for real-world genetic study design.
📝 Abstract
Longitudinal biomarkers are frequently collected in clinical studies due to their strong association with time-to-event outcomes. While considerable progress has been made in methods for jointly modeling longitudinal and survival data, comparatively little attention has been paid to statistical design considerations, particularly sample size and power calculations, in genetic studies. Yet, appropriate sample size estimation is essential for ensuring adequate power and valid inference. Genetic variants may influence event risk through both direct effects and indirect effects mediated by longitudinal biomarkers. In this paper, we derive a closed-form sample size formula for testing the overall effect of a single nucleotide polymorphism within a joint modeling framework. Simulation studies demonstrate that the proposed formula yields accurate and robust performance in finite samples. We illustrate the practical utility of our method using data from the Diabetes Control and Complications Trial.