🤖 AI Summary
This study addresses the parameter redundancy and discriminator bloat prevalent in speech bandwidth extension (BWE) models by proposing a lightweight BWE framework that integrates Swin Transformer with decision science theory. Methodologically, the generator incorporates Swin Transformer lattice interactions and linear attention mechanisms. The discriminator is reformulated through three novel lightweight architectures grounded in Conditional Value-at-Risk (CVaR), chance constraints, and multi-criteria utility theory, supplemented by differentiable barrier functions and convex combination learning to enable decision-inspired supervision. Experimental results demonstrate that the proposed generator requires only 17M parameters while achieving a 30-fold reduction in discriminator size. Furthermore, the framework surpasses existing methods in reconstruction fidelity on both English and French datasets, establishing an effective paradigm for compact yet high-performance speech enhancement.
📝 Abstract
We propose SwinDS-BWE, a decision-science-inspired bandwidth extension (BWE) model with two coupled contributions: (1) Swin Transformers with lattice-style cross-stream interaction yield locality-aware modeling with linear-in-sequence per-window attention cost and reduce the prior AP-BWE generator size 0.5x from 33M to 17M. (2) Three novel lightweight decision-science-inspired discriminators augment AP-BWE's performance: a CVaR discriminator tail-pools activations to emphasize worst high-frequency (HF) segments, a Chance-Constraint HF discriminator penalizes excessive HF energy via a differentiable barrier, and a Multi-Criteria Utility discriminator learns convex style weights over spectral criteria. SwinDS-BWE surpasses prior AP-BWE with a 30x smaller discriminator (42.3M vs. 1.36M) and higher fidelity on English and French datasets. This work shows that decision-science-inspired critics can supervise BWE to reduce discriminator size.