🤖 AI Summary
This work addresses the challenges of existing EMG-to-speech synthesis models—namely their large parameter counts, high computational complexity, and difficulty in capturing nonlinear dynamics—by proposing a lightweight architecture that integrates Samba attention with a chaos-inspired loss function. For the first time in this domain, Lyapunov Exponent Regularization (LER) and Multiscale Detrended Fluctuation Analysis (MSDFA) are incorporated, complemented by a post-hoc vocoder alignment technique. The proposed model reduces parameters by 40.79% and computational overhead by 13.33% compared to the baseline, while achieving a 2.1× improvement in Log-Spectral Distortion (LSD), a 4.7× gain in Short-Time Objective Intelligibility (STOI), and a 1.25× enhancement in Scale-Invariant Signal-to-Distortion Ratio (SI-SDR), thereby substantially improving speech reconstruction quality alongside significant model compression.
📝 Abstract
We propose a chaos-inspired new architecture for EMG-to-Speech (ETS) synthesis called CS-ETS, which combines a Samba-based encoder with two novel chaos-inspired loss functions -- Lyapunov Exponent Regularization (LER) and Multi-Scale Detrended Fluctuation Analysis (MSDFA). LER is designed based on Lyapunov exponents to capture nonlinear fluctuations and sensitivity to initial conditions. MSDFA exploits detrended fluctuation analysis to quantify fractal-like, long-range temporal chaotic correlation. CS-ETS surpasses prior work with a 40.79\% lower parameter count (32M vs 54.1M) and introduces a new Post-Vocoder Alignment approach that improves LSD by 2.1x, STOI by 4.7x, and SI-SDR by 1.25x. CS-ETS reduces computation by 13.33\% while maintaining improved performance. To the best of our knowledge, for the first time, we show how ETS can be supervised by the subtle non-linear chaotic physics with Samba attention to achieve a significantly smaller model with superior performance.