🤖 AI Summary
To address the critical bottleneck that millions of scanned paper-based electrocardiogram (ECG) recordings remain inaccessible to AI-based diagnosis, this work proposes the first robust, fully automated digitization framework for real-world clinical paper ECGs. Methodologically, it integrates four sequential stages: image preprocessing, adaptive grid detection, noise-resilient waveform tracing, and geometric correction-based signal reconstruction—effectively suppressing artifacts from noise, paper creases, stains, and perspective distortion. Our key contributions include an open-source, modular system designed for both clinical deployment and research reproducibility. Evaluated on 37,191 real-world ECG images, the framework achieves a mean signal-to-noise ratio of 19.65 dB on the Akershus dataset and consistently outperforms state-of-the-art methods across all metrics on the Emory dataset. This work establishes a high-fidelity digital signal foundation essential for scalable, AI-driven ECG interpretation.
📝 Abstract
Millions of clinical ECGs exist only as paper scans, making them unusable for modern automated diagnostics. We introduce a fully automated, modular framework that converts scanned or photographed ECGs into digital signals, suitable for both clinical and research applications. The framework is validated on 37,191 ECG images with 1,596 collected at Akershus University Hospital, where the algorithm obtains a mean signal-to-noise ratio of 19.65 dB on scanned papers with common artifacts. It is further evaluated on the Emory Paper Digitization ECG Dataset, comprising 35,595 images, including images with perspective distortion, wrinkles, and stains. The model improves on the state-of-the-art in all subcategories. The full software is released as open-source, promoting reproducibility and further development. We hope the software will contribute to unlocking retrospective ECG archives and democratize access to AI-driven diagnostics.