EHRAdapt: Adapting Pretrained Language Models to Electronic Health Records with Semantic Priors for Rare Clinical Events
This study addresses the challenges of sequence length inflation arising from serializing electronic health records and data sparsity that hinders the modeling of rare clinical events. To overcome these limitations, this work proposes an adapter that directly maps medical tuples into the embedding space of a frozen large language model. The core design introduces a “semantic prior plus low-rank residual” mechanism to jointly represent event vectors, integrated with learnable modality and temporal biases alongside a shared projection layer, thereby fine-tuning only 0.1%–0.6% of parameters. Evaluated on million-scale patient datasets, the proposed approach outperforms baseline methods and substantially improves predictive performance for long-tail rare diseases, empirically validating the complementary effectiveness of semantic priors and evidential residuals.