๐ค AI Summary
To address the limitation in bandwidth extension (BWE) that neglects the intrinsic deterministic chaos of speech, this work pioneers the integration of nonlinear dynamical modeling into the discriminative supervision process of BWE, proposing a multi-resolution chaos-aware discriminator framework. Specifically, we design two lightweight discriminators: the Multi-Scale Recurrent Discriminator (MSRD), leveraging recursive representations, and the Multi-Resolution Lyapunov Discriminator (MRLD), grounded in Lyapunov exponent estimationโboth incorporating depthwise separable convolutions and end-to-end optimization. Our approach explicitly models the chaotic dynamics underlying speech generation. It significantly reduces computational overhead while improving reconstruction quality: discriminator parameters decrease by 44ร compared to AP-BWE (22M โ 0.48M), and both objective (e.g., PESQ, STOI) and subjective (MOS) metrics show consistent improvement.
๐ Abstract
In this paper, we design two nonlinear dynamical systems-inspired discriminators -- the Multi-Scale Recurrence Discriminator (MSRD) and the Multi-Resolution Lyapunov Discriminator (MRLD) -- to extit{explicitly} model the inherent deterministic chaos of speech. MSRD is designed based on Recurrence representations to capture self-similarity dynamics. MRLD is designed based on Lyapunov exponents to capture nonlinear fluctuations and sensitivity to initial conditions. Through extensive design optimization and the use of depthwise-separable convolutions in the discriminators, our framework surpasses prior AP-BWE model with a 44x reduction in the discriminator parameter count extbf{($sim$ 22M vs $sim$ 0.48M)}. To the best of our knowledge, for the first time, this paper demonstrates how BWE can be supervised by the subtle non-linear chaotic physics of voiced sound production to achieve a significant reduction in the discriminator size.