Score
Designs and builds algorithms, models, and evaluation procedures that modify specified voice or speech attributes (such as age, gender, timbre, prosody) in recorded speech while preserving speaker identity and audio naturalness and supporting unified editing of multiple attributes. Emphasizes robustness by training and testing methods to maintain stable, successful edits in the presence of label noise and real-world acoustic noise.
Voice conversion (VC) models exhibit severe robustness deficiencies under realistic input degradations—including noise, reverberation, and adversarial perturbations—yet existing work lacks systematic analysis and quantifiable evaluation of their vulnerabilities. Method: This paper introduces the first multidimensional robustness assessment framework from an input-manipulation perspective, integrating adversarial example generation, controlled degradation injection, objective metrics (intelligibility, speaker similarity, naturalness), and subjective listening tests to quantify differential impacts of various distortions on VC output quality. Contribution/Results: Experiments reveal drastic performance degradation of state-of-the-art VC models under reverberation and adversarial attacks, confirming their reliance on non-robust acoustic features. This work fills a critical gap in systematic VC robustness research, providing empirically grounded insights and a reproducible benchmark to guide robust model design and defense strategy optimization.
This paper addresses controllable voice speaker attribute editing (e.g., age, gender) under zero-shot, non-parallel data conditions. We propose an unsupervised disentanglement framework that requires no paired samples. Methodologically, we introduce the first integration of a GAN-enhanced variational autoencoder with a two-stage acoustic conversion architecture—enabling unsupervised disentanglement and independent manipulation of speaker identity and attributes separately in the speaker embedding space and raw waveform domain. Our key contribution is the first demonstration of fine-grained, attribute-controllable editing on unpaired speech while preserving speaker identifiability and speech naturalness. Experiments show synthesized speech achieves MOS ≥ 4.1 and speaker identity preservation (cosine similarity > 0.89), significantly outperforming existing unsupervised baselines.
This work addresses the instability of speech attribute editing models caused by label noise or inconsistency in large-scale datasets. To mitigate this issue, the study introduces idempotency constraints—formally defined as \( f(f(x)) = f(x) \)—into the training of end-to-end conditional generative models, serving as an implicit regularization mechanism. By enforcing that repeated application of the editing function yields the same result as a single application, the proposed approach enhances robustness to label noise without requiring explicit noise modeling. Experiments demonstrate significant improvements in editing success rates on both synthetically controlled noisy data and the real-world GLOBE dataset, while simultaneously better preserving speaker identity characteristics compared to existing methods.
This work addresses the limitations of existing speech editing methods that operate in the acoustic space, where content and style are often entangled, leading to generation instability and boundary artifacts that hinder perceptually seamless text-driven editing. To overcome these challenges, the authors propose an “edit content, preserve acoustics” framework that decouples editing into the semantic space and employs a Flow Matching decoder to reconstruct acoustic features. Additionally, a self-consistency reward mechanism is introduced, leveraging a pre-trained text-to-speech (TTS) model as an implicit evaluator to enforce context-aware alignment. Experimental results demonstrate that the proposed approach significantly outperforms current autoregressive and non-autoregressive baselines in intelligibility, robustness, and perceptual quality, achieving high-fidelity, seamlessly edited speech.
Existing text-to-speech (TTS) models struggle to directly modify reference timbres and achieve segment-level local expressiveness control. To address these limitations, this work proposes EDICT, a framework that pioneers the construction of shared timbre anchors within the codec token space to unify global timbre editing with local expressiveness control. Specifically, the method generates edited acoustic references to anchor target timbres and integrates segment-wise instructions with dynamic KV cache reconstruction techniques, effectively balancing instruction following, speaker consistency, and transition quality. Experimental results demonstrate that EDICT significantly improves timbre editing performance and overall quality on benchmarks such as TimbreEdit-Bench.
In real-world acoustic scenes, users struggle to manipulate unseparated mixed sound sources. To address this, we propose the first end-to-end, text-driven framework for real-time sound field editing, enabling direct, joint manipulation of multiple concurrent sources—such as “reduce air-conditioner noise and enhance speech”—via natural language instructions, without explicit source separation. Our method integrates large language model–based semantic parsing with a differentiable spectrogram decomposition–filtering–reconstruction architecture, supporting open-vocabulary, zero-shot editing. It is trained on a newly curated dataset comprising 160 hours of audio and 100,000 audio–text pairs. Experiments demonstrate significant improvements in source extraction, suppression, and level control: +3.2 dB in SI-SNR and +0.11 in STOI, while maintaining strong robustness and generalization across complex mixtures with 2–5 overlapping sources.
Existing speech editing methods are largely confined to word-level content modification and treat speaker identity, emotion, and linguistic content as separate tasks, lacking fine-grained control and unified editing capabilities. This work proposes UniSAE, a unified framework that, for the first time, enables composable editing of speaker characteristics, emotional attributes, and multi-granular speech content within a single architecture. UniSAE introduces Discrete Phoneme Posterior Graphs (DPPGs) as sub-phonemic representations, integrating an autoregressive content transformer with a diffusion-based acoustic decoder, alongside disentangled speaker and emotion embeddings. The framework supports high-precision joint editing ranging from sub-phonemic to word-level units, significantly enhancing editing flexibility and control accuracy while preserving naturalness in the synthesized speech.
该论文介绍了一个开源基础模型AuK,通过自然语言指令和音频上下文统一语音生成和编辑,使用多任务训练及后训练策略优化性能。
This work addresses a critical gap in the evaluation of controllable speech generation systems, which typically focus on target prompt alignment while neglecting the stability of non-target attributes. The study systematically reveals, for the first time, that adjusting a target attribute often induces unintended shifts in other acoustic or speaker characteristics. To investigate this phenomenon, the authors construct an audit dataset comprising 5,940 samples and perform controlled pairwise evaluations using multidimensional metrics spanning acoustics, prosody, content, and speaker traits. They further propose VoDER-Cal, a training-free, inference-time candidate re-ranking method that significantly enhances attribute preservation. In a three-candidate setting, VoDER-Cal improves joint success rate from 4.8% to 14% and reduces non-target deviation from 0.344 to 0.276, demonstrating its effectiveness in mitigating unintended perturbations during targeted voice editing.
This work addresses the challenge of imprecise control in existing speech editing methods under natural language instructions, which often suffer from semantic ambiguity in specifying edit types, parameters, and target regions. To overcome this, the authors propose a structured editing interface grounded in transcribed text, employing XML-style tags to explicitly denote operation types and anchor them to specific transcript spans or boundaries, thereby constructing a semantic timeline that circumvents the need for explicit time alignment. Building upon this framework, they enhance the continuous autoregressive model dots.tts to support four composable editing dimensions—lexical content, emotion, prosody, and pauses—while preserving contextual integrity in unedited segments. The contributions include the first structured instruction framework for speech editing, a task-oriented data curation pipeline, and doteBench, the first bilingual benchmark for precise evaluation. Experiments demonstrate state-of-the-art instruction-following accuracy and local fidelity across five editing tasks in doteBench, with audio quality comparable to leading open-source systems and no significant degradation in zero-shot TTS error rates or speaker similarity relative to the base model.
This work addresses the lack of a unified benchmark for speech editing evaluation, which hinders the simultaneous assessment of multi-attribute manipulation and preservation of irrelevant characteristics. To bridge this gap, we propose the first unified evaluation framework supporting multi-attribute editing, composite instructions, and bilingual (Chinese–English) inputs. We construct a benchmark dataset encompassing seven atomic and composite editing tasks and introduce an anchor-based contrastive evaluation protocol with three fine-grained metrics: target success, preservation success, and joint success. Combining human and automatic evaluations, we conduct a comprehensive assessment of leading speech foundation models and specialized systems, revealing critical limitations: imbalanced performance across editing dimensions, superior efficacy of closed-source over open-source models, and notably low joint success rates on composite tasks.