ASCENDgpt: A Phenotype-Aware Transformer Model for Cardiovascular Risk Prediction from Electronic Health Records

๐Ÿ“… 2025-08-31
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the challenge of multi-endpoint cardiovascular risk prediction from longitudinal electronic health records (EHRs). To this end, we propose a phenotype-aware Transformer model. Methodologically, we design a clinically interpretable phenotype tokenization scheme that compresses tens of thousands of diagnosis codes into 176 semantically cohesive phenotype tokensโ€”reducing vocabulary size by 77.9%. The model integrates masked language modeling pretraining with time-to-event fine-tuning to jointly predict five time-to-event endpoints: myocardial infarction, stroke, major adverse cardiovascular events (MACE), cardiovascular death, and all-cause death. On an independent test set, the model achieves a mean C-index of 0.816 (0.842 for cardiovascular death), significantly improving predictive accuracy, clinical interpretability, and computational efficiency. This work establishes a novel paradigm for dynamic, EHR-driven cardiovascular risk assessment.

Technology Category

Machine Learning: Hardware-aware MLComputer Vision: Interpretability, Explainability, and TransparencyNatural Language Processing: Interpretability, Analysis, and Evaluation of NLP Models

Application Category

User Modeling, Personalization and Recommendation: Explainable and interpretable methods for personalizationSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsEconomics, Online Markets and Human Computation: Trust and reliance of crowd workers and data experts on GenAI
๐Ÿ“ Abstract
We present ASCENDgpt, a transformer-based model specifically designed for cardiovascular risk prediction from longitudinal electronic health records (EHRs). Our approach introduces a novel phenotype-aware tokenization scheme that maps 47,155 raw ICD codes to 176 clinically meaningful phenotype tokens, achieving 99.6% consolidation of diagnosis codes while preserving semantic information. This phenotype mapping contributes to a total vocabulary of 10,442 tokens - a 77.9% reduction when compared with using raw ICD codes directly. We pretrain ASCENDgpt on sequences derived from 19402 unique individuals using a masked language modeling objective, then fine-tune for time-to-event prediction of five cardiovascular outcomes: myocardial infarction (MI), stroke, major adverse cardiovascular events (MACE), cardiovascular death, and all-cause mortality. Our model achieves excellent discrimination on the held-out test set with an average C-index of 0.816, demonstrating strong performance across all outcomes (MI: 0.792, stroke: 0.824, MACE: 0.800, cardiovascular death: 0.842, all-cause mortality: 0.824). The phenotype-based approach enables clinically interpretable predictions while maintaining computational efficiency. Our work demonstrates the effectiveness of domain-specific tokenization and pretraining for EHR-based risk prediction tasks.
Problem

Research questions and friction points this paper is trying to address.

Predicting cardiovascular risk from longitudinal electronic health records
Mapping raw ICD codes to clinically meaningful phenotype tokens
Achieving interpretable predictions while maintaining computational efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Phenotype-aware tokenization mapping raw ICD codes
Transformer model pretrained with masked language modeling
Fine-tuned for time-to-event cardiovascular outcome prediction
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.