🤖 AI Summary
This study addresses the lack of systematic evaluation regarding the transferability and adaptability of foundation models for electromyography (EMG) signals. To this end, we construct a unified benchmark by integrating 20 datasets to evaluate nine pretrained models under linear probing, full fine-tuning, and few-shot adaptation paradigms, providing the first systematic quantification of pretraining benefits. Our findings reveal that pretraining does not universally confer advantages, whereas full fine-tuning plays a critical role in performance enhancement. Furthermore, we identify consistent cross-limb patterns and demonstrate that 5-shot adaptation only partially recovers performance. This work offers essential empirical guidance for the selection and deployment of EMG foundation models.
📝 Abstract
Foundation models (FMs) are increasingly being developed for general time series and physiological signals, yet their transferability to downstream physiological tasks remains poorly understood. This question is particularly challenging for electromyography (EMG), where signal distributions vary substantially across users, sensing configurations, acquisition hardware, and downstream tasks. We introduce EMG-FM-Bench, a systematic benchmark for studying foundation-model transfer and adaptation on EMG. EMG-FM-Bench unifies 20 public datasets with over 1 million EMG segments and evaluates nine pretrained foundation models across four questions: how pretrained models perform when frozen or fully fine-tuned, how much pretraining helps compared with training the same model from scratch, how well models generalize to new users with limited labeled data, and how performance changes across different EMG tasks. Across the benchmark, linear probing provides useful information about pretrained representations, but full fine-tuning can substantially change downstream EMG performance. Comparing each pretrained model with the same model trained from scratch shows that the benefit of pretraining varies substantially across models and is not universal. Performance decreases when models are evaluated on new users, while five-shot adaptation improves macro-F1 in 70.2% of evaluated model-dataset combinations but recovers only part of the lost performance. Model performance is highly consistent between upper- and lower-limb classification and remains strongly correlated with continuous EMG-to-text decoding. Together, these results provide a systematic view of when pretrained time-series models transfer effectively to EMG and how their performance depends on fine-tuning, user variation, and downstream task.