🤖 AI Summary
This study addresses the challenges of noise interference in electronic health record discharge summaries and the tendency of generic summarization to omit information critical for clinical prediction. We propose a reinforcement learning framework that utilizes downstream clinical prediction feedback as a reward signal, directly aligning large language model summarizers with task-specific objectives. Furthermore, a longitudinal encoder generates soft prompts for multimodal fusion, precisely extracting patient-specific textual evidence complementary to structured encodings. Evaluated on the MIMIC dataset, our approach significantly improves readmission prediction and medication recommendation performance, outperforming strong existing baselines. These results validate the effectiveness of task-oriented summarization in enhancing predictive modeling from clinical narratives.
📝 Abstract
Unstructured discharge notes in Electronic Health Records (EHRs) often carry signal complementary to structured medical codes, holding patient-specific evidence that standardized cohort-level codes alone cannot capture. However, this evidence in notes is frequently buried in lengthy, noisy text that is not intentionally written with any specific clinical prediction in mind. Summarization is an obvious mitigation, but generic summaries, tuned for fluency rather than the outcome, routinely omit decisive evidence while retaining plausible but uninformative detail. To this end, we propose RASPER, a Reward-Aligned Summarizer for Prediction in EHR, that optimizes note summarization directly against the downstream clinical task. RASPER employs a tunable LLM-based summarizer to extract task-relevant evidence from discharge notes and trains it via reinforcement learning from prediction feedback, using a reward derived from the downstream predictor's loss. To ground the summarizer, a longitudinal encoder converts structured codes into soft prompts that incorporate each patient's clinical context into note summarization. By rewarding the quality of the resulting multimodal prediction, RASPER encourages the summarizer to retain patient-specific evidence that complements, rather than duplicates, information captured by structured codes. RASPER consistently outperforms strong baselines on both readmission prediction and medication recommendation across MIMIC-III and MIMIC-IV.