SEmoEdit: Probing and Harnessing the Editability of Pre-trained Speech Flows

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the reliance of existing speech emotion editing methods on extensive task-specific training and their inherent instability. We propose the first training-free framework that formulates emotion editing as dynamic velocity transport between source and target states. By leveraging the generative dynamics of flow matching and hybrid TTS models, our approach achieves architecture-aware editing without parameter updates through diagnostic generation trajectories, enabling robust emotion replacement, erasure, and continuous interpolation directly within pretrained models. This work is the first to reveal that pretrained speech flows harbor rich latent emotion editing capabilities. Furthermore, we construct SEmoEditBench, a benchmark dataset comprising 600 cases. Experiments demonstrate that our method comprehensively outperforms existing training-based and activation-steering approaches across state-of-the-art models, validating its broad applicability within pretrained speech flows.
📝 Abstract
Existing training-based speech emotion editing methods often require substantial task-specific training and can be unstable. This motivates us to investigate whether the pretrained generative dynamics of large-scale text-to-speech (TTS) models can be directly manipulated for training-free emotion editing. To answer this question, we probe the editability of pretrained flow-matching and hybrid TTS models by constructing a controlled test set and systematically diagnosing editing effects along the generative trajectory. Our analysis reveals that pretrained TTS models are substantially editable in emotion, but such editability is architecture- and trajectory-dependent and can be disrupted by early flow-matching steps, while cross-speaker emotion transport carries additional acoustic attributes beyond emotion. To address these limitations, we propose SEmoEdit, the first training-free framework that formulates emotion editing as dynamic velocity transport between source and target emotions, enabling robust, flow-based speech emotion editing directly within pretrained TTS models. SEmoEdit unifies three core operations: emotion replacement, emotion erasure, and continuous emotion interpolation, requiring neither parameter updates nor task-specific optimization. To systematically evaluate these capabilities, we introduce SEmoEditBench, a dataset comprising 600 editing cases, and conduct extensive experiments across state-of-the-art (SOTA) models and backbones. Our results show that SEmoEdit is highly effective and broadly applicable, outperforming existing training-based and activation-steering methods. Ultimately, this work reveals that pretrained speech flows possess rich, latent emotion-editing capabilities, providing useful guidance for real applications. Code, benchmark, and Audio samples are available at https://github.com/imxtx/SEmoEdit.
Problem

Research questions and friction points this paper is trying to address.

speech emotion editing
training-free
pre-trained TTS models
flow-matching
editability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Training-free emotion editing
Flow-matching TTS
Dynamic velocity transport
Speech generation
SEmoEditBench
🔎 Similar Papers
No similar papers found.