🤖 AI Summary
Existing ECG foundation models predominantly rely on single-source interpretation reports, limiting their capacity to fully exploit the broad diagnostic signals embedded in electrocardiographic waveforms. To address this limitation, this work proposes MS-ECG-FM, a framework that introduces a multi-source contrastive learning mechanism to align ECGs with multimodal clinical texts—including echocardiography, radiology, and discharge summaries—during cross-modal pretraining, thereby overcoming the constraints of single-supervision paradigms. By fusing complementary information across modalities, the proposed approach significantly enhances the comprehensiveness and robustness of representation learning. Extensive evaluations demonstrate that MS-ECG-FM consistently outperforms existing methods across multiple detection benchmarks, exhibiting particularly strong performance under reduced-lead configurations and achieving efficient generalization across diverse clinical tasks.
📝 Abstract
Electrocardiography (ECG) records the electrical activity of the heart, aiding diagnosis by detecting abnormalities in cardiac function. ECG foundation models have demonstrated promising results, but are limited by a reliance on ECG interpretation reports as their sole supervision. Because interpretation reports only capture the subset of waveform information routinely recognized by clinicians, this constrains representation learning to overlook the broader diagnostic signals present in ECG. We introduce a new ECG foundation model --- MS-ECG-FM --- that is trained through contrastive alignment to multiple distinct clinical note types, including ECG, echocardiography, radiology, and discharge reports. We evaluate MS-ECG-FM on an extended set of ECG detection benchmarks, showing that it comprehensively outperforms existing methods on the full span of conditions that ECG can detect, including in reduced-lead configurations. Different reports improve representations for different diagnostic domains, while multi-source alignment captures their complementary information and produces consistently strong representations across clinically diverse tasks.