Score
Designs and implements signal-processing pipelines and tools that clean and condition raw electrocardiogram (ECG) recordings for analysis, including baseline-wander removal, bandpass and notch filtering, resampling and detrending, interpolation for missing samples, artifact and noise detection/correction, beat segmentation and alignment, and per-lead selection or aggregation. Also develops and applies signal-quality metrics and validation procedures to produce time-series suitable for feature extraction, automated interpretation, or other downstream algorithms.
This study addresses the lack of consensus on preprocessing strategies for multi-label ECG-based cardiac disease classification. We systematically evaluate the impact of downsampling, normalization, and bandpass filtering on three state-of-the-art time-series classifiers—InceptionTime, TS-Transformer, and ROCKET—across three benchmark ECG datasets. Results show that 50 Hz sampling achieves performance comparable to 500 Hz, reducing model parameter count and training time by ~90%; min-max normalization slightly degrades accuracy, while IIR/FIR bandpass filtering yields no significant improvement; Z-score normalization demonstrates robustness. Crucially, we empirically refute both the “preprocessing-irrelevance” and “blind-preprocessing-effectiveness” hypotheses, providing the first evidence that preprocessing must be task-adaptive: sensitivity varies significantly across cardiac disease subtypes and model architectures. Our findings establish a reproducible, lightweight, and task-driven preprocessing paradigm for ECG-based intelligent diagnosis.
Existing ECG analysis studies often overlook the electrophysiological characteristics of ECG signals and clinical application requirements, leading to inadequate evaluation frameworks. To address this, we propose ECG-Bench—the first comprehensive, multi-task benchmark for ECG time-series analysis—covering four clinically relevant downstream tasks: rhythm classification, anomaly detection, lesion localization, and risk prediction. We introduce novel evaluation metrics tailored to ECG spectral properties and clinical interpretability. Furthermore, we design ECGFormer, a lightweight, physiology-aware temporal modeling architecture. Leveraging a large-scale pretraining benchmarking framework, systematic evaluations demonstrate that our metrics improve assessment accuracy by +12.3% on average, while ECGFormer achieves a mean F1-score of 94.7% across six mainstream datasets (e.g., MIT-BIH), outperforming state-of-the-art models by 3.1%. This work establishes a standardized evaluation paradigm and delivers a high-performance foundational model for intelligent ECG analysis.
Digital conversion of paper-based single-lead electrocardiogram (ECG) images suffers from degraded accuracy due to signal overlap—a long-standing, under-addressed challenge. Method: We propose an end-to-end robust digitization framework: (1) a customized data-augmented U-Net for precise segmentation of the primary ECG waveform; and (2) an adaptive grid detection module that faithfully maps the binary segmentation mask into a time-series signal. Contribution/Results: To our knowledge, this is the first method explicitly designed for overlapping ECG traces, achieving superior generalization across diverse formats and multi-scale images. Quantitatively, on overlapping samples, mean squared error (MSE) drops to 0.0029 (84% reduction over baseline) and Pearson’s correlation coefficient ρ reaches 0.9641; on non-overlapping samples, MSE = 0.0010 and ρ = 0.9644; segmentation achieves an IoU of 0.87. The implementation is publicly available.
Existing ECG noise detection models suffer from limited generalizability, as most studies evaluate performance on single, homogeneous datasets, failing to assess robustness across diverse noise types and acquisition conditions. Method: This paper proposes a generalizable noise detection framework grounded in heart rate variability (HRV) features. It integrates time–frequency domain HRV feature engineering with machine learning classifiers and adopts an AUPRC-centric evaluation paradigm coupled with multi-dataset cross-validation. Contribution/Results: We conduct the first systematic cross-dataset evaluation across four heterogeneous public ECG datasets, assessing generalization against motion artifacts, electromyographic interference, and other noise classes. The framework achieves an average accuracy of 90.2% and an AUPRC of 0.912 on unseen datasets—significantly outperforming single-dataset baselines. Results demonstrate the strong robustness and broad applicability of HRV-driven approaches in cross-source, multi-noise-type scenarios.
Real-world ECG analysis faces significant challenges, including strong data heterogeneity, high noise levels, substantial inter-population variability, and complex rhythm-event associations. To address these, we propose AnyECG—the first foundation model tailored for multi-source heterogeneous ECG data. Our method introduces three key innovations: (1) an ECG Tokenizer that encodes continuous, noisy signals into discrete, compact, and clinically interpretable local rhythm tokens; (2) a proxy-task-driven discretized representation learning framework coupled with rhythm-pattern-aware autoregressive pretraining; and (3) joint self-supervised learning across diverse devices and clinical scenarios. Evaluated on four tasks—abnormality detection, arrhythmia classification, lead imputation, and ultra-long ECG analysis—AnyECG consistently surpasses state-of-the-art methods, demonstrating marked improvements in generalization and robustness under realistic, noisy, and heterogeneous conditions.
This study addresses the challenge of denoising canine electrocardiogram (ECG) signals, which are highly susceptible to complex noise sources such as respiration, electromyographic interference, and lead artifacts. Conventional denoising approaches often fail to simultaneously suppress noise and preserve diagnostically critical waveform morphologies. To overcome this limitation, the authors propose an end-to-end deep learning denoising model based on an autoencoder architecture that reconstructs clean ECG signals as a preprocessing step, thereby significantly enhancing the accuracy of downstream AI-driven waveform segmentation. The method innovatively integrates denoising with the optimization of the subsequent task, demonstrating robust performance across diverse noise conditions. It effectively retains morphological features of diagnostic value and consistently outperforms existing techniques in both signal fidelity and segmentation accuracy.
This study addresses the lack of systematic evaluation of automated electrocardiogram (ECG) interval measurement methods in large-scale real-world clinical cohorts. The authors propose an end-to-end waveform parsing system evaluated on 10,646 twelve-lead ECGs to assess the accuracy of PR, QRS, and QT/QTc intervals. Their approach introduces a novel Fast Fourier Convolution ResNet (FFCResNet) with register tokens to jointly model local temporal and global spectral features, complemented by ECG-specific data augmentation. Results demonstrate mean absolute errors of 17.5 ms for QT interval, 14.8 ms for QRS duration, and 0.8 beats per minute for ventricular rate. P-, QRS-, and T-wave segmentation achieved Dice scores of 95.5%–98.2%. Algorithmic errors approached inter-observer variability in sinus rhythm but increased significantly during supraventricular tachycardia, thereby delineating the first clear boundary of clinical applicability.
This work proposes ECG-LENS, an end-to-end framework for multi-lead electrocardiogram (ECG) report generation aimed at alleviating clinician workload and improving diagnostic efficiency. The approach integrates a lead-aware encoder with global dependency modeling, augmented by a clinical terminology–enhanced textual prompting mechanism and an ECG-specific report preprocessing strategy to enable diagnosis-aware, context-guided text generation. Additionally, the authors introduce F1-ECGBERT, a BERT-based evaluation metric tailored to ECG report assessment. Evaluated on the PTB-XL and MIMIC-IV-ECG datasets, the model substantially outperforms existing methods, achieving relative improvements of 4.0% in METEOR, 6.3% in ROUGE-L, and 11.5% in F1-ECGBERT.
This work addresses the limitations of existing ECG classification systems, whose optimization typically relies on manual failure analysis and struggles with coarse-grained, aggregate metrics that obscure specific error sources. To overcome this, we propose RecursiveECG, a novel framework that formalizes clinical diagnostic criteria into executable measurement functions. By integrating waveform data, derived measurements, and model predictions, RecursiveECG enables evidence-driven failure auditing. Leveraging a large language model as an offline designer, the framework recursively refines the classifier based on concrete failure cases. Evaluated on PTB-XL, Georgia, and CPSC2018 benchmarks, RecursiveECG achieves an average relative performance gain of 10.0% over strong baselines. Notably, it incurs no LLM inference overhead at deployment and supports auditable, traceable model revisions.