neural audio synthesis

Designing and training neural or codec-based synthesis pipelines to generate realistic synthetic audio (speech, sound effects, codec-transformed signals) for dataset creation and evaluation across speaker, language, and event classes.

neuralaudiosynthesis

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models

Sep 07, 2025
YY
Yi Yuan
🏛️ University of Surrey | Seed Group | ByteDance Inc.

Current text-to-audio (T2A) models struggle to achieve precise, fine-grained control over acoustic attributes, hindering personalized audio generation. To address this, we propose the first framework tailored for customized T2A synthesis, leveraging reference-audio-guided diffusion modeling integrated with large language models and dual-modality alignment training—enabling joint modeling of semantic fidelity and target acoustic characteristics (e.g., timbre, rhythm, event structure). Our method extracts and faithfully reproduces speaker- or style-specific acoustic features from only a few reference samples. We introduce two dedicated datasets, including the CTTA benchmark with real-world scene annotations. Experiments demonstrate that our model significantly outperforms state-of-the-art methods on customized generation tasks, achieving superior audio quality, semantic alignment, and acoustic feature fidelity—while maintaining competitive performance on general-purpose T2A benchmarks.

Generating customized audio events from user-provided reference samplesPrecise control of fine-grained acoustic characteristics in audio generationSemantic alignment between generated audio and input text prompts

To address the challenges of subjective quality assessment and the inefficiency of purely data-driven models in speech/audio coding, this paper proposes a tightly integrated hybrid neural coding framework that synergistically combines model-driven and data-driven paradigms. Methodologically, it introduces a novel multi-level hybrid architecture that deeply couples psychoacoustic-weighted loss, customized time-frequency domain prediction (TF-Codec/MDCTNet), an LPCNet-based backbone, and a neural post-processing module, trained end-to-end via an autoencoder paradigm. The core contribution lies in systematically bridging the performance gap between classical signal modeling and end-to-end deep learning. Experimental results demonstrate that, at ultra-low bitrates of 1.6–3.2 kbps, the proposed method achieves a P.808 MOS gain of ≥0.5 over baselines, yielding subjective audio quality approaching that of wideband codecs, while increasing computational overhead by less than 15%.

Efficiency ImprovementNeural Voice and Audio CodingQuality Evaluation

ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs For Audio, Music, and Speech

Sep 24, 2024
JS
Jiatong Shi
🏛️ Carnegie Mellon University | Renmin University of China | Nanyang Technological University | Tokyo Metropolitan University | University of Chicago | National Taiwan University

To address the challenges of incomparable cross-task performance and narrow evaluation dimensions in neural audio, music, and speech coding, this paper introduces the first open-source unified training and evaluation platform. Methodologically, it proposes ESPnet-Codec—an integrated codec framework—and VERSA—a standalone evaluation toolkit—supporting mainstream models (e.g., SoundStream, EnCodec, DAC) with discrete quantization, residual vector quantization (RVQ), and multi-scale adversarial training. It enables fully automated assessment across 20 objective audio metrics and seamless integration with six ESPnet downstream tasks. The key contributions are: (1) establishing the first cross-modal neural codec benchmark; (2) significantly improving training efficiency and interoperability for downstream applications (e.g., TTS, music generation); and (3) achieving state-of-the-art performance on both objective metrics and subjective MOS scores, with full reproducibility.

Challenges in fair comparisons across applicationsComprehensive evaluation of codec performanceNeural codecs for audio, music, and speech

Contrastive Learning from Synthetic Audio Doppelgangers

Jun 09, 2024
MC
Manuel Cherep
🏛️ Massachusetts Institute of Technology

Audio representation learning heavily relies on large-scale real-world recordings, while manual annotation and data augmentation struggle to capture the full diversity of physical acoustics. Method: We propose a synthetic-driven contrastive learning framework that requires no real audio data. It employs differentiable and stochastic sound synthesizers to generate physically consistent synthetic “twin” positive pairs online via causally interpretable parameter perturbations (e.g., timbre, pitch, envelope), thereby constructing high-diversity contrastive tasks. Contribution/Results: We introduce the first positive-pair construction paradigm grounded in causal perturbations of synthesizer parameters; require only a single interpretable hyperparameter and zero real-data storage; and achieve, for the first time, synthetic-data-only models that surpass real-data baselines on ESC-50, UrbanSound8K, and SpeechCommands. This significantly reduces data dependency and storage overhead, establishing a new paradigm for low-resource audio representation learning.

Existing audio transformations lack true real-world sound diversity.Learning robust audio representations requires large real-world datasets.Synthetic audio doppelgängers provide rich contrastive learning information.

Synthio: Augmenting Small-Scale Audio Classification Datasets with Synthetic Data

Oct 02, 2024
SG
Sreyan Ghosh
🏛️ NVIDIA | University of Maryland

To address data scarcity in low-resource audio classification, this paper proposes a novel data augmentation framework integrating text-to-audio (T2A) diffusion modeling, preference optimization via proximal policy optimization (PPO), and large language model (LLM)-driven iterative caption generation. The method employs PPO to align synthesized audio with target acoustic characteristics, while the LLM dynamically generates and refines semantically diverse, high-quality captions—jointly enhancing acoustic fidelity and semantic richness. Distinct from prior work, this is the first study to holistically unify T2A diffusion, human-feedback-based preference optimization, and LLM-guided iterative captioning. Evaluated across 10 benchmark datasets under four low-resource settings, the framework achieves substantial performance gains (+0.1%–39%) using only a weakly supervised AudioSet-pretrained T2A model, consistently outperforming state-of-the-art baselines.

Ensuring acoustic consistency and diversity in synthetic dataGenerating synthetic audio that matches real-world diversityImproving audio classification with limited labeled data

Latest Papers

What's happening recently
View more

Training end-to-end text-to-speech (TTS) models with purely synthetic data remains underexplored, particularly regarding feasibility, robustness, and controllability compared to real speech data. Method: This study systematically evaluates FastSpeech 2- and VITS-based TTS models trained exclusively on synthetic speech, conducting ablation experiments by controlling textual richness, speaker diversity, environmental noise level, and speaking style. Evaluation integrates MOS, WER, CMOS, and subjective listening tests. Contribution/Results: To our knowledge, this is the first empirical demonstration that synthetic-data-only training achieves a MOS of 4.12—significantly surpassing the real-data baseline (3.78) at equivalent scale. The synthetic-trained models exhibit 27% higher robustness to accent and noise, and 31% improved cross-speaker generalization similarity. Key findings identify high text/speaker diversity and low environmental noise as primary drivers of robustness, while standard speaking style accelerates convergence. These results establish a theoretically grounded, cost-effective paradigm for controllable, high-quality TTS data curation.

Assessing synthetic data's potential to outperform real data trainingExploring factors like speaker diversity and noise affecting model performanceInvestigating feasibility of purely synthetic data for TTS training

This work addresses the absence of evaluation benchmarks capable of precisely matching real-world audio recordings with their conditionally generated synthetic counterparts. The authors introduce Doppelganger, a benchmark comprising 10,420 real-synthetic sound effect pairs and seven categories of controlled stimuli, and formally define, for the first time, a cross-domain sound matching task spanning real and synthetic audio. Using a contrastive embedding model trained on paired data, the approach achieves 80% source identification accuracy on unseen sound events—substantially outperforming a class-label baseline (61%) and random chance (0.03%). Additional experiments reveal that human listeners misidentify synthetic sounds as real with 29% probability, whereas generator-specific detectors can nearly perfectly discriminate between the two. This study underscores the critical role of paired supervision in learning effective cross-domain audio representations.

audio benchmarkaudio matchingreal-synthetic boundary

This work addresses the ambiguity in labeling resynthesized audio generated by neural audio codecs—a class of models that combine compression and synthesis capabilities—within the context of voice spoofing detection. The study presents the first systematic analysis of this labeling challenge, introducing an extended version of the ASVspoof 5 dataset and proposing multiple annotation strategies tailored to resynthesized audio. A unified evaluation framework is designed to assess the impact of different labeling approaches on anti-spoofing systems, leveraging resynthesis techniques that integrate neural codecs with vocoders. Experimental results demonstrate that the choice of annotation strategy significantly influences detection performance, offering critical insights for future dataset construction and evaluation protocols in audio deepfake detection research.

audio deepfake detectionlabeling ambiguityneural audio codecs

This work addresses the challenge of unified generation of speech, sound effects, and music—a task requiring high fidelity, end-to-end training, effective contextual modeling, and variable-length synthesis, which existing approaches struggle to balance. The authors propose AudioCALM, a framework that extends autoregressive language modeling to continuous audio latent spaces by integrating flow matching (predicting rectified-flow velocities) with a block-causal AR-Flow attention mechanism, enabling high-quality generation of arbitrary duration. A novel asymmetric multimodal mixture-of-experts architecture (A-MoME) and a unified conditioning interface are introduced: speech utilizes dedicated residual experts, while sound effects and music share a common backbone, incurring no additional inference cost. Experiments demonstrate that AudioCALM matches or surpasses specialized models across all three modalities and significantly outperforms current unified audio generation methods.

autoregressive modelingmultimodal asymmetrytext-audio alignment

Codec2Vec: Self-Supervised Speech Representation Learning Using Neural Speech Codecs

Nov 20, 2025
WT
Wei-Cheng Tseng
🏛️ University of Texas at Austin

To address the substantial storage overhead, low transmission efficiency, and heightened privacy risks inherent in continuous speech representations, this paper proposes Codec2Vec—the first self-supervised speech representation learning framework entirely based on discrete acoustic units from neural audio codecs. Codec2Vec employs a masked discrete unit prediction objective, jointly optimizing multiple targets to learn robust and compact speech representations. On the SUPERB benchmark, it achieves performance comparable to continuous-input models while reducing storage requirements by up to 16.5× and accelerating training by 2.3×. Crucially, its discrete tokenization enables native on-device data anonymization, significantly enhancing privacy preservation and system scalability. The core contribution lies in pioneering the direct modeling of discrete speech codes as the fundamental units for self-supervised learning—establishing a new paradigm for efficient, secure, and lightweight speech representation learning.

Developing self-supervised speech representation learning using discrete codec unitsExploring masked prediction strategies for efficient speech processingReducing storage and training requirements while maintaining competitive performance

Hot Scholars

YM

Yuki Mitsufuji

Distinguished Engineer, Sony
Machine LearningAudioSource SeparationMusic Technology
HY

Hung-yi Lee

National Taiwan University
deep learningspoken language understandingspeech processing
ZZ

Zhou Zhao

Zhejiang University
Machine LearningData MiningMultimedia Computing
SW

Shinji Watanabe

Carnegie Mellon University
Speech recognitionSpeech processingSpeech enhancementSpeech translation
ZW

Zhizheng Wu

The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen), Mel Lab
Spoken Language ProcessingDeepFake detectionMusic Processing