🤖 AI Summary
This work addresses the challenges of data heterogeneity, label scarcity, and the absence of a unified representation in electromyography (EMG) signals across users, devices, and tasks. To this end, we propose AEMG—the first large-scale self-supervised representation learning framework tailored for EMG. Our approach uniquely models neuromuscular dynamics as a “physiological language,” leveraging a Neuromuscular Contraction Tokenizer (NCT) to discretize continuous EMG signals into “words” and “sentences,” thereby constructing the largest cross-device EMG vocabulary to date. AEMG introduces a unified self-supervised pretraining paradigm that accommodates arbitrary channel topologies and sampling rates. Experiments demonstrate that AEMG improves accuracy by 5.79–9.25% under zero-shot leave-one-user-out evaluation and achieves over 90% few-shot adaptation performance using only 5% of target-user data.
📝 Abstract
A fundamental role in decoding human motor intent and enabling intuitive human-computer interaction is played by electromyography (EMG). However, its generalization capability across subjects, devices, and tasks remains substantially limited by data heterogeneity, label scarcity, and the lack of a unified representational framework. To bridge this gap, we propose Any Electromyography (AEMG), the first large-scale, self-supervised representation learning framework for EMG. AEMG reconceptualizes neuromuscular dynamics linguistically, utilizing a novel Neuromuscular Contraction Tokenizer (NCT) to translate discrete muscle contractions into structural words and temporal activation patterns into coherent sentences. Furthermore, we compile the largest cross-device EMG signal vocabulary to date, enabling seamless transfer across arbitrary channel topologies and sampling rates. Experiments demonstrate that AEMG improves the zero-shot leave-one-subject-out (LOSO) accuracy by 5.79-9.25% compared to six state-of-the-art baselines, and achieves more than 90% few-shot adaptation performance with only 5% of target user data. Our work has proposed the concept of EMG signals as a cross-device physiological language, learned their grammar from massive amounts of data, and laid the groundwork for a single-training, universally applicable EMG foundation model.