Detecting gene-environment interactions to guide personalized intervention: boosting distributional regression for polygenic scores

📅 2025-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Current polygenic risk score (PRS) methods model only phenotypic means, neglecting the genetic regulation of phenotypic variance and thus failing to capture gene–environment interactions (G×E). To address this, we propose snpboostLSS—the first sparse PRS algorithm that integrates gradient boosting into the Gaussian location–scale (LSS) framework, enabling joint estimation of polygenic effects on both phenotypic mean and variance. Leveraging iterative gradient boosting coupled with batch-wise SNP screening, snpboostLSS achieves efficient distributional regression in high-dimensional genomic settings. Applied to UK Biobank data, it robustly detects G×E signals: between statin use and polygenic risk for LDL cholesterol, and between BMI polygenic scores and physical activity/sedentary behavior. These findings advance mechanistic understanding of context-dependent genetic effects and establish a novel paradigm for context-aware, personalized intervention strategies.

Technology Category

Machine Learning: Ensemble MethodsSearch and Optimization: Mixed Discrete/Continuous SearchNatural Language Processing: Discourse, Pragmatics & Argument Mining

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsUser Modeling, Personalization and Recommendation: Explainable and interpretable methods for personalizationSearch and Retrieval-Augmented AI: Personalized, context-aware and across-device search
📝 Abstract
Polygenic risk scores can be used to model the individual genetic liability for human traits. Current methods primarily focus on modeling the mean of a phenotype neglecting the variance. However, genetic variants associated with phenotypic variance can provide important insights to gene-environment interaction studies. To overcome this, we propose snpboostlss, a cyclical gradient boosting algorithm for a Gaussian location-scale model to jointly derive sparse polygenic models for both the mean and the variance of a quantitative phenotype. To improve computational efficiency on high-dimensional and large-scale genotype data (large n and large p), we only consider a batch of most relevant variants in each boosting step. We investigate the effect of statins therapy (the environmental factor) on low-density lipoprotein in the UK Biobank cohort using the new snpboostlss algorithm. We are able to verify the interaction between statins usage and the polygenic risk scores for phenotypic variance in both cross sectional and longitudinal analyses. Particularly, following the spirit of target trial emulation, we observe that the treatment effect of statins is more substantial in people with higher polygenic risk scores for phenotypic variance, indicating gene-environment interaction. When applying to body mass index, the newly constructed polygenic risk scores for variance show significant interaction with physical activity and sedentary behavior. Therefore, the polygenic risk scores for phenotypic variance derived by snpboostlss have potential to identify individuals that could benefit more from environmental changes (e.g. medical intervention and lifestyle changes).
Problem

Research questions and friction points this paper is trying to address.

Modeling both mean and variance of phenotypes using polygenic scores
Detecting gene-environment interactions for personalized intervention strategies
Improving computational efficiency for high-dimensional genotype data analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cyclical gradient boosting for joint mean and variance modeling
Batch processing of relevant variants for computational efficiency
Polygenic variance scores to identify gene-environment interactions
🔎 Similar Papers
No similar papers found.
Q
Qiong Wu
Institute for Medical Biometry and Statistics, Marburg University, Germany
H
Hannah Klinkhammer
Institute for Genomic Statistics and Bioinformatics, University of Bonn, Germany
K
Kiran Kunwar
Center for Human Genetics, Marburg University, Germany
C
Christian Staerk
IUF-Leibniz Research Institute for Environmental Medicine, Düsseldorf, Germany
Carlo Maj
Carlo Maj
Head of Bioinformatics at Center for Human Genetics, University of Marburg
Statistical Genetics
A
Andreas Mayr
Institute for Medical Biometry and Statistics, Marburg University, Germany