speech

Designs, builds, or evaluates systems that capture, represent, model, transform, or generate human spoken audio (speech/voice, 语音), including components for recognition, synthesis, enhancement, separation, speaker identification/verification, and voice conversion or detection.

speech

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-2.1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$195K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

What Do Humans Hear When Interacting? Experiments on Selective Listening for Evaluating ASR of Spoken Dialogue Systems

Aug 06, 2025
KM
Kiyotada Mori
🏛️ Nara Institute of Science and Technology | Guardian Robot Project | RIKEN | National Institute of Informatics | Tokyo Institute of Technology

This study addresses the gap between automatic speech recognition (ASR) and human auditory cognition in spoken dialogue systems (SDS), specifically investigating how humans perform selective listening during dialogue and what recognition capabilities ASR must acquire to approach human performance. Using experimental psychology paradigms, we quantify human information selection preferences in natural conversations via manual transcription analysis, dialogue response generation tasks, and attention pattern modeling. We propose a novel, cognition-grounded ASR evaluation framework—first to operationalize selective listening as measurable cognitive metrics. Experiments reveal that humans consistently ignore redundant acoustic segments and prioritize semantically critical units (e.g., intent verbs, entity nouns), whereas state-of-the-art ASR systems exhibit systematic deficits in capturing such units. Our work establishes a cognitively informed benchmark for ASR evaluation, advancing the field from lexical accuracy toward semantic relevance.

Compares human and ASR transcription for response generationInvestigates human selective listening in dialogue interactionsProposes new ASR evaluation method using human listening patterns

This study addresses the misuse of AI-generated speech and associated detection challenges by presenting the first systematic integration of generation and detection technologies across the full pipeline. Through constructing a technical taxonomy, curating benchmark resources, and establishing an open challenge framework, this work develops a comprehensive knowledge map of the field. The research not only clarifies technological evolution and critical bottlenecks but also delineates a future roadmap tailored to speech-specific characteristics. By providing systematic theoretical support and practical guidance for building robust speech security defenses, this survey fills a significant gap in existing literature regarding holistic, end-to-end perspectives on AI speech synthesis and forensics.

AI-generated voicedeepfake audiodisinformation

Emerging AI-based voice attacks increasingly involve mixed audio—comprising authentic speech, fully synthetic speech, voice clones, and hybrid combinations—posing novel security threats to voice authentication systems. Method: To address this, we introduce the first comprehensive benchmark dataset covering all four audio categories and propose a novel hybrid audio detection framework based on fine-tuned Audio Spectrogram Transformer (AST). Unlike conventional binary classification approaches, our method pioneers a hybrid acoustic pattern modeling paradigm, integrating spectrogram-based representation learning, transfer learning, and controllable hybrid audio synthesis with precise annotation. Contribution/Results: Experiments demonstrate that our approach achieves 97% accuracy on hybrid audio detection—significantly outperforming existing baselines. This work fills critical gaps in both data resources and modeling methodologies for hybrid voice detection, establishing a reproducible benchmark and an effective technical pathway to enhance the robustness of speaker verification systems against sophisticated voice spoofing attacks.

Addressing security risks in voice authentication from advanced cloningDetecting hybrid audio mixing human and AI-generated speech segmentsImproving detection accuracy for complex mixed acoustic patterns

STAR: Speech-to-Audio Generation via Representation Learning

Sep 21, 2025
ZX
Zeyu Xie
🏛️ Peking University | Shanghai Jiao Tong University

Existing speech-to-audio generation systems predominantly adopt cascaded architectures, suffering from low inference efficiency and severe error propagation. To address these limitations, we propose STAR—the first end-to-end speech-to-audio generation framework that directly utilizes raw speech as the interactive control signal for audio synthesis. STAR employs deep representation learning to extract sound events and scene semantics from speech, introduces a bridging network to map speech representations to multimodal audio features, and adopts a two-stage training strategy to jointly optimize representation learning and audio synthesis. Experiments demonstrate that STAR reduces speech processing latency by 76.9% compared to cascaded baselines, while significantly outperforming them in audio fidelity, event accuracy, and scene consistency. These results validate the feasibility and superiority of end-to-end, speech-driven audio generation.

Addressing error propagation in cascaded speech processing systemsDeveloping first end-to-end speech-to-audio generation frameworkExtracting sound event semantics directly from raw speech signals

In real-world acoustic scenes, users struggle to manipulate unseparated mixed sound sources. To address this, we propose the first end-to-end, text-driven framework for real-time sound field editing, enabling direct, joint manipulation of multiple concurrent sources—such as “reduce air-conditioner noise and enhance speech”—via natural language instructions, without explicit source separation. Our method integrates large language model–based semantic parsing with a differentiable spectrogram decomposition–filtering–reconstruction architecture, supporting open-vocabulary, zero-shot editing. It is trained on a newly curated dataset comprising 160 hours of audio and 100,000 audio–text pairs. Experiments demonstrate significant improvements in source extraction, suppression, and level control: +3.2 dB in SI-SNR and +0.11 in STOI, while maintaining strong robustness and generalization across complex mixtures with 2–5 overlapping sources.

Control sound sources in mixtures via text instructionsEnhance auditory experience with semantic text filtersRemix multiple sounds simultaneously without separation

Latest Papers

What's happening recently
View more

This study investigates whether individual dimensions in the representations of self-supervised speech models (specifically WavLM) encode distinct speaker-related acoustic attributes, such as pitch, gender, intensity, noise level, and the second formant. By applying principal component analysis (PCA) to disentangle model features, the authors systematically identify independent dimensions that exhibit strong correlations with these acoustic properties, establishing for the first time a clear correspondence between specific latent dimensions and interpretable speaker characteristics. Further experiments demonstrate that manipulating these dominant dimensions enables effective control over the associated speaker attributes in speech synthesis, thereby confirming both their controllability and practical utility in downstream applications.

dimension analysisself-supervised speech featuresspeaker characteristics

This study addresses the lack of a unified quantitative evaluation framework for assessing the disentanglement of speaker identity and prosody in speech content representations. To this end, we construct a generative model relying solely on a single representation to systematically compare self-supervised learning (SSL) features and supervised tokens across content, identity, and prosody dimensions. Our findings reveal that disentanglement efficacy is governed by the interplay between training objectives and information capacity, rather than being determined exclusively by supervisory signals. Furthermore, we identify two distinct representational paradigms: high-fidelity reconstruction and strong disentanglement. We demonstrate that, under constrained capacity, supervised representations can effectively isolate speaker identity. These insights provide a novel theoretical foundation for advancing speech representation learning.

disentanglementgenerative frameworkspeaker identity

This work addresses the ambiguity in labeling resynthesized audio generated by neural audio codecs—a class of models that combine compression and synthesis capabilities—within the context of voice spoofing detection. The study presents the first systematic analysis of this labeling challenge, introducing an extended version of the ASVspoof 5 dataset and proposing multiple annotation strategies tailored to resynthesized audio. A unified evaluation framework is designed to assess the impact of different labeling approaches on anti-spoofing systems, leveraging resynthesis techniques that integrate neural codecs with vocoders. Experimental results demonstrate that the choice of annotation strategy significantly influences detection performance, offering critical insights for future dataset construction and evaluation protocols in audio deepfake detection research.

audio deepfake detectionlabeling ambiguityneural audio codecs

This work proposes CookVoice, a unified non-autoregressive framework that enables multi-modal and multi-task human voice generation—including text-to-speech, text-to-singing, voice cloning, conversion, and editing—within a single model, addressing the limitations of existing systems that are often task-specific, autoregressive, and lacking in fine-grained control and inference efficiency. By decomposing voice into content, prosody, and style components and incorporating a frame-level alignment mechanism with an ordinary differential equation (ODE) solver, CookVoice achieves flexible control signal mapping and high-quality synthesis. With only 43.51 million parameters, the model generates both speech and singing in as few as four ODE steps, significantly enhancing controllability, generalization, and inference speed.

controllabilityinference efficiencymulti-modal

This work addresses the limitations in current non-human voice conversion research, which has been hindered by the absence of publicly available, structured datasets and an overreliance on natural human speech. To bridge this gap, we present the first systematically constructed and open-sourced dataset encompassing both human and animal vocalizations, augmented with professionally designed audio effects. The dataset explicitly disentangles timbre and style dimensions and incorporates a structured train-test split strategy, enabling rigorous evaluation of models’ controllable generalization capabilities across both seen and unseen timbres and styles. By providing a high-quality benchmark and reproducible foundation, this resource significantly advances research in non-human voice conversion.

audio datasetdesigned vocalizationsnon-human voices