Is Self-Pretraining really useful to improve diagnosis in medical Time Series?

📅 2026-08-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether self-supervised pre-training (SPT) can effectively enhance the diagnostic performance of Transformer models on multimodal, multivariate, and univariate medical time-series data, particularly in data-scarce clinical settings. The work proposes a general-purpose approach that requires no task-specific architectural modifications and incorporates four masking strategies to facilitate representation learning across both modalities and temporal dimensions. Experimental results across three medical time-series tasks demonstrate that SPT improves classification accuracy by 0–6 percentage points, with deeper models exhibiting more pronounced gains. Notably, this is the first study to validate the efficacy of SPT on univariate medical time-series data, demonstrating its strong generalizability, scalability, and robustness under limited-data conditions.
📝 Abstract
Inspired by recent evidence that transformer architectures benefit from Self-PreTraining (SPT) on long-context benchmarks, we investigate whether similar gains extend to multimodal, multivariate, and even simple univariate medical time series. Our objective is to assess the impact of SPT on the performance and scalability of transformer-based models across diverse medical applications, particularly under limited data conditions. We evaluate transformer architectures on three representative medical time-series tasks: rehabilitation robotics (Camargo dataset), stress detection (Non-EEG Stress), and Parkinson's disease detection (Gait Parkinson's Disease). Models are trained either from scratch or through SPT using four masking-based objectives designed to promote temporal and cross-modal representation learning, and we systematically vary model depth to examine how capacity interacts with pre-training benefits. Across datasets and configurations, SPT consistently improves classification accuracy by 0-6 percentage points depending on masking strategy, dataset and architecture, with gains observed not only in multivariate settings but also when models are restricted to simple univariate inputs. The improvements increase for deeper models that can better exploit the enriched temporal representations learned during pre-training. These findings indicate that SPT is a simple and general strategy that enhances transformer performance on medical time-series tasks without requiring task-specific architectural changes, supporting its potential to improve robustness and accuracy in data-limited clinical settings.
Problem

Research questions and friction points this paper is trying to address.

Self-PreTraining
medical time series
diagnosis
transformer
limited data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-PreTraining
medical time series
transformer
masking-based pretraining
data-limited learning
🔎 Similar Papers
No similar papers found.
O
Omar Coser
Unit of Artificial Intelligence & Computer Systems, Università Campus Bio-Medico di Roma, Via Álvaro del Portillo, 21, Rome, 00128, Italy; Unit of Advanced Robotics and Human-Centered Technologies, Università Campus Bio-Medico di Roma, Via Álvaro del Portillo, 21, Rome, 00128, Italy
Antonio Orvieto
Antonio Orvieto
ELLIS Institute Tübingen, Max Planck Institute for Intelligent Systems
Deep LearningMachine LearningOptimizationDifferential EquationsNumerical Analysis
Paolo Soda
Paolo Soda
Professor of AI, Università Campus Bio-Medico di Roma, Italy
Artificial intelligencemachine learninghealthcaremedical imaging
L
Loredana Zollo
Unit of Advanced Robotics and Human-Centered Technologies, Università Campus Bio-Medico di Roma, Via Álvaro del Portillo, 21, Rome, 00128, Italy