🤖 AI Summary
This study addresses the lack of low-bitrate streaming neural codecs for electromyography (EMG) signals and their limited cross-task generalization. To this end, we propose a streaming neural codec-decoder tailored for EMG. Methodologically, we introduce a novel architecture integrating a causal Transformer with residual vector quantization (RVQ), enabling efficient transmission and discrete tokenized representations, while multi-dataset joint training enhances generalizability. The model supports real-time streaming inference at 50 Hz and facilitates seamless integration with large language models. It achieves state-of-the-art performance across downstream tasks such as typing and pose estimation, requiring only 0.482 ms of computation per frame. The source code has been made publicly available.
📝 Abstract
Neural codecs encode continuous signals into compact sequences of discrete tokens, providing an interface for efficient transmission, storage, and token-based sequence modeling. This paradigm has been widely adopted in modern speech and audio frameworks; however, the biosignal domain still lacks a neural codec designed specifically for low-bitrate streaming and generalization across diverse downstream tasks. We present MyoCodec, a streaming neural codec designed for electromyography (EMG). Inspired by recent neural audio codecs, MyoCodec combines causal Transformers with residual vector quantization to encode continuous EMG signals into different levels of EMG representations spanning from continuous latent features to discrete tokens operating at 50 Hz. Trained on twelve public EMG datasets, MyoCodec achieves favorable performance in both intrinsic codec quality and representative downstream tasks, including typing (emg2qwerty), hand-pose (emg2pose), speech decoding (emg2speech), and speech-to-EMG synthesis (speech2emg). Across these tasks, MyoCodec exhibits strong performance against prior models while providing a compact and causal EMG representation. During streaming inference, it requires compute time of only 0.482 ms for each 20 ms frame, enabling real-time streaming. Also, the discrete token representation provided by MyoCodec has the potential to support integration into language-model based approaches, creating a path toward LLM-based interactive systems, where tokenized EMG representations are directly processed into such language or speech models. Code and model weights are released.