🤖 AI Summary
This study addresses the lack of clinical alignment in PPG foundation models and the difficulty large language models face in interpreting waveform physiological information by introducing the first PPG-language model family. Methodologically, it proposes a three-level segment-event-encounter automatic EHR annotation pipeline alongside a two-stage cross-modal alignment framework. By integrating contrastive learning, waveform-conditioned generation, and temporal-aware aggregation, the model undergoes multimodal pretraining on approximately 73,000 hours of data, achieving deep fusion of waveform signals with medical knowledge. Experimental results demonstrate that the proposed approach significantly enhances retrieval accuracy and descriptive factuality on benchmarks such as MC-MED, while outperforming existing baselines across multiple clinical prediction tasks.
📝 Abstract
Photoplethysmography (PPG) is widely recorded by clinical monitors and consumer wearables, providing a scalable source of continuous physiological information. These recordings offer an opportunity for physiological assessment at scale, but realizing this potential requires models to learn from both signal-derived physiological supervision and broader clinical context captured in electronic health records (EHRs). This involves aligning information spanning local observations, care events, and entire visits with PPG representations at corresponding temporal scales. However, existing PPG foundation models primarily rely on task-specific prediction heads, while the medical knowledge of large language models does not necessarily translate into waveform understanding. To bridge this gap, we introduce PPG-LM, the first PPG-language model family to learn physiological representations from both signal-derived supervision and broader clinical context captured in EHRs. To construct clinically grounded captions, we develop an automatic captioning pipeline that generates segment-, event-, and visit-level descriptions from signal measurements and structured EHR records. We then learn from these pairs through a two-stage framework that first establishes segment-language correspondence through contrastive learning and waveform-conditioned captioning, then extends alignment to events and visits through time-aware aggregation and temporal statement matching. Pretrained on approximately 73k hours of PPG, PPG-LM supports language-based recognition, cross-modal retrieval, and segment captioning. Experiments on MC-MED, MIMIC-III, and VitalDB show improved retrieval and caption factuality over language-model baselines and gains over PPG and time-series foundation models on multiple clinical prediction tasks.