π€ AI Summary
This work addresses the challenge of effectively leveraging multivariate time-series dataβsuch as 12-lead electrocardiograms (ECGs)βin medical domains where labeled data are scarce. The authors propose the Event Reconstruction Joint-Embedding Predictive Architecture (ER-JEPA), which introduces a novel Hierarchical JEPA (H-JEPA) model. This architecture employs a two-stage hierarchical design: it first models local temporal segments and then treats the resulting representations as a univariate sequence for global modeling, all within a Vision Transformer backbone trained via self-supervised learning. Inspired by clinical ECG interpretation workflows, the approach enables multi-level abstract representation learning. Pretrained on only approximately 180,000 ten-second ECG recordings, the model achieves state-of-the-art performance on the ST-MEM benchmark while maintaining computational efficiency and low resource consumption.
π Abstract
Data analysis in the medical domain often encounters scenarios involving a limited target dataset and a large, unannotated dataset with a general distribution. Under such circumstances, self-supervised learning (SSL) methods are highly effective for utilizing large datasets, making them a popular choice for electrocardiogram (ECG) analysis. This work presents the Event Reconstruction Joint-Embedding Predictive Architecture (ER-JEPA), a lightweight SSL framework for multivariate time series, whose name and two-fold hierarchical structure are inspired by the diagnostic approach of cardiologists. At its core, ER-JEPA features: (1) a two-stage structure that constructs representations for each time interval and subsequently processes these representations as a univariate time series, (2) the hierarchical integration of two Joint-Embedding Predictive Architectures (JEPAs), and (3) a Vision Transformer (ViT) backbone. The structural concatenation of two JEPAs categorizes the model as a Hierarchical JEPA (H-JEPA), designed to encode multiple levels of abstract representations for enhanced prediction on complex tasks. This study reports a successful application of H-JEPA to 12-lead ECG data as a multivariate time series alongside an analysis of the sensitivity of hierarchical representation during the pretraining stage. Pretrained on approximately 180,000 10-second recordings, the model achieves state-of-the-art downstream performance on the ST-MEM benchmark, with rapid computation and minimal resource usage.