🤖 AI Summary
This study addresses the challenges of catastrophic forgetting and insufficient zero-shot generalization in speech enhancement models during continual multi-domain learning. To this end, it proposes the first framework integrating generative language models with domain-incremental learning for speech enhancement. By leveraging a pretrained language model architecture alongside lightweight low-rank adaptation (LoRA), the approach enables incremental adaptation to novel acoustic domains. Crucially, this paradigm employs LoRA to simultaneously acquire capabilities in new domains while effectively preserving previously learned knowledge. Extensive experiments conducted across four heterogeneous datasets demonstrate that the proposed model successfully adapts to unseen domains while significantly mitigating performance degradation on prior ones, thereby establishing an effective solution for robust, domain-continual speech enhancement.
📝 Abstract
We propose a domain-incremental learning framework for generative speech enhancement (SE) that learns from a sequence of datasets or domains recorded under diverse acoustic conditions. Fine-tuning a pretrained model on continuously evolving domains leads to catastrophic forgetting of previously acquired knowledge, while zero-shot generalization often fails to adequately adapt to unseen domains. To address these challenges, we first develop a novel language model-based generative SE model that we then use as a pretrained backbone and incrementally adapt it to acoustically mismatched domains using lightweight domain-specific Low-Rank Adaptation. The proposed framework enables the model to acquire enhancement capabilities for new domains while preserving performance on previously learned domains. Evaluated on four heterogeneous speech datasets, our approach effectively adapts to new domains without forgetting previously learned domains.