🤖 AI Summary
This study addresses the challenges of data heterogeneity and modeling bottlenecks in predicting opioid overdose risk from longitudinal ICD diagnosis histories. To this end, it proposes OverdoseMoE, a mixture-of-experts ensemble framework that integrates models of varying scales through an optimized complementary weighting strategy. Leveraging Mamba and Qwen architectures, the approach combines continuous pre-training with task-specific fine-tuning of large language models to achieve diagnosis-specific adaptation. The proposed model attains an AUPRC of 25.17 for 180-day overdose risk prediction, with a positive predictive value of 25.38% among the top 5% highest-risk individuals. Furthermore, cross-cohort validation demonstrates strong robustness, indicating that OverdoseMoE provides precise and reliable decision support for early clinical intervention.
📝 Abstract
Opioid overdose remains a major clinical and public health burden, highlighting the need for scalable approaches to identify patients at high risk. Here, we investigate diagnosis-specific adaptation for 180-day opioid overdose risk prediction from patients' preceding one-year longitudinal ICD histories. We develop OODMAMBA and OODQWEN through continued pretraining on longitudinal diagnostic sequences followed by task-specific fine-tuning. Building on the stronger Qwen-based predictors, we further propose OVERDOSEMOE, a multi-expert framework that integrates models of different scales using complementary expert-weighting strategies. Diagnosis-specific adaptation consistently improved predictive performance over general-purpose language-model baselines, with OODQWEN achieving an AUPRC of 24.47 and an AUROC of 68.56. OVERDOSEMOE further improved discrimination and precision, achieving an AUPRC of 25.17 and an AUROC of 69.49 while outperforming the strongest single-model baselines. Among patients ranked in the top 5% of predicted risk, OVERDOSEMOE identified substantially enriched overdose risk, achieving a PPV of 25.38% while retaining meaningful recall. Evaluation on an independent MIMIC-IV cohort further demonstrated cross-cohort robustness, with complementary weighting strategies showing advantages across different performance measures. These findings demonstrate that diagnosis-specific language-model adaptation combined with multi-expert integration can improve opioid overdose risk stratification and support more robust prediction across heterogeneous electronic health record populations.