A Semantic Information-based Hierarchical Speech Enhancement Method Using Factorized Codec and Diffusion Model

📅 2025-05-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing speech enhancement methods predominantly rely on time-frequency masking or spectral estimation, neglecting the intrinsic coupling between semantic content and acoustic details—leading to insufficient robustness under challenging conditions (e.g., low SNR, strong reverberation) and degraded downstream TTS performance. To address this, we propose a semantics-informed hierarchical modeling framework: (1) a dual-stream semantic–acoustic architecture that explicitly disentangles these two representations for the first time in speech enhancement; and (2) a factorized encoder–decoder coupled with a conditional diffusion model to enable coarse-to-fine joint time-frequency reconstruction. Experiments demonstrate state-of-the-art performance across objective metrics (PESQ, STOI) and end-to-end TTS quality (intelligibility and naturalness), with particularly pronounced gains under low SNR (−5 dB to 5 dB) and highly reverberant conditions.

Technology Category

Natural Language Processing: SpeechMachine Learning: Semi-Supervised LearningComputer Vision: Diffusion Models for Vision

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Large language models for searchUser Modeling, Personalization and Recommendation: User modeling and simulation for interactive and conversational systems
📝 Abstract
Most current speech enhancement (SE) methods recover clean speech from noisy inputs by directly estimating time-frequency masks or spectrums. However, these approaches often neglect the distinct attributes, such as semantic content and acoustic details, inherent in speech signals, which can hinder performance in downstream tasks. Moreover, their effectiveness tends to degrade in complex acoustic environments. To overcome these challenges, we propose a novel, semantic information-based, step-by-step factorized SE method using factorized codec and diffusion model. Unlike traditional SE methods, our hierarchical modeling of semantic and acoustic attributes enables more robust clean speech recovery, particularly in challenging acoustic scenarios. Moreover, this method offers further advantages for downstream TTS tasks. Experimental results demonstrate that our algorithm not only outperforms SOTA baselines in terms of speech quality but also enhances TTS performance in noisy environments.
Problem

Research questions and friction points this paper is trying to address.

Recover clean speech while preserving semantic and acoustic attributes
Improve performance in complex acoustic environments
Enhance downstream TTS tasks in noisy conditions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hierarchical modeling of semantic and acoustic attributes
Uses factorized codec for step-by-step speech enhancement
Incorporates diffusion model for robust clean speech recovery
🔎 Similar Papers
Y
Yang Xiang
Hithink Research, Zhejiang, China; Electrical Engineering Department, Tsinghua University, Beijing, China
C
Canan Huang
Hithink Research, Zhejiang, China
D
Desheng Hu
Hithink Research, Zhejiang, China
Jingguang Tian
Jingguang Tian
Midea AI Research Institute, Shanghai, China
speech LLMaudio understanding
X
Xinhui Hu
Hithink Research, Zhejiang, China
C
Chao Zhang
Electrical Engineering Department, Tsinghua University, Beijing, China