EHRAdapt: Adapting Pretrained Language Models to Electronic Health Records with Semantic Priors for Rare Clinical Events

๐Ÿ“… 2026-09-27
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the challenges of sequence length inflation arising from serializing electronic health records and data sparsity that hinders the modeling of rare clinical events. To overcome these limitations, this work proposes an adapter that directly maps medical tuples into the embedding space of a frozen large language model. The core design introduces a โ€œsemantic prior plus low-rank residualโ€ mechanism to jointly represent event vectors, integrated with learnable modality and temporal biases alongside a shared projection layer, thereby fine-tuning only 0.1%โ€“0.6% of parameters. Evaluated on million-scale patient datasets, the proposed approach outperforms baseline methods and substantially improves predictive performance for long-tail rare diseases, empirically validating the complementary effectiveness of semantic priors and evidential residuals.
๐Ÿ“ Abstract
Electronic health records (EHRs) encode clinical histories as (time, modality, code) tuples, whereas pretrained language models expect text tokens. Serializing them as text inflates sequence length and redundantly encodes structure. We introduce EHRAdapt, an adapter that maps tuples directly into a frozen language model's embedding space. Modality receives a learned embedding, time gaps enter through learned attention biases, and event codes receive dedicated vectors. Learning event vectors is the central challenge: clinical vocabularies are long-tailed, leaving rare events too few observations for reliable estimates. EHRAdapt therefore represents each event vector as the sum of a semantic prior and an evidence residual. The prior is a frozen embedding of the event's clinical description from a biomedical language model trained on clinical ontologies, mapped into the model's input space by a shared learned projection, so it supplies clinical meaning even when observations are scarce. The residual, a learned low-rank event-specific correction, refines it as evidence accumulates. We run continued pretraining on about 4 million patients'records with three frozen LLM backbones (OLMo2 1B, Llama3.2 1B, and OLMo2 7B), training only the adapter (0.1--0.6% of all parameters). The full adapter outperforms all ablations in held-out next-event prediction on every backbone. Removing the semantic pathway hurts rare events over ten times more than the most frequent ones, whereas removing the residual hurts overall prediction but improves it for the rarest events. On reportable infectious-disease and syndromic downstream classification tasks, EHRAdapt outperforms text-based LLM and count-based baselines, and both pathways improve rare-disease discrimination. The two pathways therefore play complementary roles, visible only when results are broken down by event frequency rather than averaged.
Problem

Research questions and friction points this paper is trying to address.

Electronic Health Records
Pretrained Language Models
Rare Clinical Events
Long-tailed Distribution
Representation Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

EHRAdapt
Semantic Priors
Parameter-Efficient Adaptation
Rare Clinical Events
Low-Rank Residual
๐Ÿ”Ž Similar Papers
No similar papers found.