Lyapunov-Inspired LyRIC Activation and GLARE Attention in Chaos-Guided State Space Modeling for EMG-To-Speech (ETS) Synthesis

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the neglect of chaotic dynamics in electrolaryngeal-to-speech (ETS) synthesis and the over-smoothing problem inherent in Transformers by pioneering the integration of nonlinear chaotic physics into neural networks. Methodologically, it introduces the LyRIC activation function and Lyapunov exponent regularization to model acoustic dynamics, combined with a DFA loss, Mamba state-space models, and GLARE attention to construct a compact, highly interpretable architecture. Experimental results demonstrate that the proposed model improves objective intelligibility by 4.69× and spectral reconstruction by 2.08× while reducing parameter count by 73.49%, establishing a new state-of-the-art for real-time ETS synthesis.
📝 Abstract
Electromyography-to-Speech (ETS) synthesis is typically a non-linear, chaotic dynamical system. However, no prior work has studied the chaotic behavior of ETS synthesis to date. Yet, prior works strictly rely on standard reconstruction metrics with parameter-heavy transformers that systematically over-smooth natural acoustic dynamics. To close this gap, for the first time, we propose a chaos-inspired Lyapunov-derived activation function (LyRIC) with two novel chaotic loss functions, Lyapunov Exponent Regularization and Multi-Scale Detrended Fluctuation Analysis, to explicitly capture the deterministic chaos of human phonation. In addition, we introduce a compressed novel encoder, GLAME, which synergizes global Mamba state-space modeling with localized GLARE attention. We comprehensively perform frame-level acoustic evaluation in a multilingual and multi-speaker setup using English and Mandarin datasets. The proposed system outperforms the established baseline with a 4.69x increase in objective intelligibility (STOI: 0.61 vs. 0.13) and a 2.08x improvement in spectral reconstruction (LSD: 1.08 vs. 2.25). Importantly, this improvement is achieved with 73.49% fewer parameters (14.34M vs. 54.10M), establishing a new baseline for ETS synthesis. To the best of our knowledge, this is the first work demonstrating that integrating non-linear chaotic physics into neural networks yields superior yet compact inductive biases for real-time ETS synthesis.
Problem

Research questions and friction points this paper is trying to address.

EMG-to-Speech synthesis
chaotic dynamics
over-smoothing
parameter-heavy models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Lyapunov-inspired activation
Chaos-guided state space modeling
GLARE attention
EMG-to-speech synthesis
Mamba encoder
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Sajid Fardin Dipto
Department of Cyber Security Engineering, George Mason University, Fairfax, VA, USA
Tarikul Islam Tamiti
Tarikul Islam Tamiti
Graduate Research Assistant at George Mason University
Generative AINatural Language Processing
L
Luke Baja-Ricketts
Department of Cyber Security Engineering, George Mason University, Fairfax, VA, USA
D
David Vergano
Department of Cyber Security Engineering, George Mason University, Fairfax, VA, USA
Anomadarshi Barua
Anomadarshi Barua
Assistant Professor @ Cyber Security Engineering, George Mason University
Cyber Physical SystemLow -Power System DesignSub-Nyquist SamplingComputer Micro Architecture