NLDSI-BWE: Non Linear Dynamical Systems-Inspired Multi Resolution Discriminators for Speech Bandwidth Extension

๐Ÿ“… 2025-10-01
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
To address the limitation in bandwidth extension (BWE) that neglects the intrinsic deterministic chaos of speech, this work pioneers the integration of nonlinear dynamical modeling into the discriminative supervision process of BWE, proposing a multi-resolution chaos-aware discriminator framework. Specifically, we design two lightweight discriminators: the Multi-Scale Recurrent Discriminator (MSRD), leveraging recursive representations, and the Multi-Resolution Lyapunov Discriminator (MRLD), grounded in Lyapunov exponent estimationโ€”both incorporating depthwise separable convolutions and end-to-end optimization. Our approach explicitly models the chaotic dynamics underlying speech generation. It significantly reduces computational overhead while improving reconstruction quality: discriminator parameters decrease by 44ร— compared to AP-BWE (22M โ†’ 0.48M), and both objective (e.g., PESQ, STOI) and subjective (MOS) metrics show consistent improvement.

Technology Category

Machine Learning: Multimodal LearningNatural Language Processing: GenerationCognitive Modeling & Cognitive Systems: Computational Creativity

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: User modeling and simulation for interactive and conversational systemsResponsible Web: Machine-in-the-loop, human agency and autonomy
๐Ÿ“ Abstract
In this paper, we design two nonlinear dynamical systems-inspired discriminators -- the Multi-Scale Recurrence Discriminator (MSRD) and the Multi-Resolution Lyapunov Discriminator (MRLD) -- to extit{explicitly} model the inherent deterministic chaos of speech. MSRD is designed based on Recurrence representations to capture self-similarity dynamics. MRLD is designed based on Lyapunov exponents to capture nonlinear fluctuations and sensitivity to initial conditions. Through extensive design optimization and the use of depthwise-separable convolutions in the discriminators, our framework surpasses prior AP-BWE model with a 44x reduction in the discriminator parameter count extbf{($sim$ 22M vs $sim$ 0.48M)}. To the best of our knowledge, for the first time, this paper demonstrates how BWE can be supervised by the subtle non-linear chaotic physics of voiced sound production to achieve a significant reduction in the discriminator size.
Problem

Research questions and friction points this paper is trying to address.

Designing chaotic dynamics-inspired discriminators for speech bandwidth extension
Modeling deterministic chaos in speech using recurrence and Lyapunov exponents
Reducing discriminator parameters while maintaining bandwidth extension performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Nonlinear dynamical systems-inspired discriminators model speech chaos
Multi-scale recurrence captures self-similarity dynamics
Multi-resolution Lyapunov exponents capture nonlinear fluctuations
๐Ÿ”Ž Similar Papers
No similar papers found.