speech and audio

Designs, builds, or analyzes algorithms, models, and systems that process, transform, generate, or extract information from audio and speech signals; this includes tasks such as feature extraction and acoustic modeling, speech recognition and synthesis, source separation and enhancement, speaker diarization and identification, audio event detection and classification, and objective or perceptual evaluation of audio quality.

speechandaudio

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$228K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Large Audio-Language Models (LALMs) suffer from pervasive hallucination—particularly in sound event existence judgment, temporal relation reasoning, and sound source attribute identification—undermining their real-world reliability. To address this, we introduce the first systematic, multi-task benchmark for audio hallucination evaluation, comprising three fine-grained tasks: existence verification, temporal ordering, and attribute alignment. We further propose a novel multi-round chain-of-reasoning framework that integrates audio-language joint prompting with stepwise thinking mechanisms to explicitly model foundational audio semantic logic. Experiments demonstrate that our approach significantly mitigates spurious generation, temporal misalignment, and source-attribute mismatches, achieving an average accuracy improvement of 23.6% across all three tasks. This work is the first to systematically identify and enhance LALMs’ robustness at the fundamental audio semantic level.

Accuracy IssuesAudio Language ModelsReal-world Applications

Computer Audition: From Task-Specific Machine Learning to Foundation Models

Jul 22, 2024
AT
Andreas Triantafyllopoulos
🏛️ Technical University of Munich | Tampere University | University of Augsburg | Imperial College | Munich Center for Machine Learning | Munich Data Science Institute

To address the longstanding reliance on single-task models and the lack of generalizable representations in computer audition, this paper proposes a systematic framework for constructing Auditory Foundation Models (AFMs). Methodologically, it establishes the core paradigm of AFMs for the first time, integrating unified multi-task modeling, cross-modal (audio–text) aligned representation learning, and instruction-driven human–machine interaction. Technically, the framework encompasses large-scale audio–text contrastive pretraining, multi-task prompt tuning, and self-supervised audio modeling. Experiments demonstrate that the proposed AFM achieves substantial performance gains across 10+ downstream tasks—including automatic speech recognition, sound source separation, and environmental sound classification—while enabling zero-shot transfer and open-domain speech understanding. This work advances computer audition toward generality, multi-task synergy, and natural human–machine interaction.

Consolidate multiple audio tasks into single foundation modelsLeverage cross-modal knowledge for general-purpose audio understandingTransition from task-specific models to auditory foundation models

In real-world acoustic scenes, users struggle to manipulate unseparated mixed sound sources. To address this, we propose the first end-to-end, text-driven framework for real-time sound field editing, enabling direct, joint manipulation of multiple concurrent sources—such as “reduce air-conditioner noise and enhance speech”—via natural language instructions, without explicit source separation. Our method integrates large language model–based semantic parsing with a differentiable spectrogram decomposition–filtering–reconstruction architecture, supporting open-vocabulary, zero-shot editing. It is trained on a newly curated dataset comprising 160 hours of audio and 100,000 audio–text pairs. Experiments demonstrate significant improvements in source extraction, suppression, and level control: +3.2 dB in SI-SNR and +0.11 in STOI, while maintaining strong robustness and generalization across complex mixtures with 2–5 overlapping sources.

Control sound sources in mixtures via text instructionsEnhance auditory experience with semantic text filtersRemix multiple sounds simultaneously without separation

Audio-Language Datasets of Scenes and Events: A Survey

Jul 09, 2024
GW
Gijs Wijngaard
🏛️ Maastricht University

This study systematically evaluates 69 audio-language datasets available as of September 2024, revealing pervasive issues including acoustic class imbalance, multi-source duplication, linguistic homogeneity (dominant English bias), restricted accessibility, and latent societal biases. Methodologically, we innovatively integrate PCA-based cross-dataset embedding variance analysis, CLAP-guided detection of modality leakage, joint acoustic–textual distribution modeling, and open governance practices to quantitatively identify systemic biases—particularly in widely used sources such as YouTube and Freesound. As a key contribution, we release an open resource library comprising over two million samples and propose a comprehensive Audio-Language Modeling (ALM) data curation roadmap that explicitly balances diversity, robustness, and fairness. This work establishes an empirically grounded, reproducible methodology for dataset development, directly supporting improved generalization capabilities of multimodal models.

Audio-lingual Model TrainingData Bias and LimitationsDataset Analysis

USED: Universal Speaker Extraction and Diarization

Sep 19, 2023
JA
Junyi Ao
🏛️ The Chinese University of Hong Kong | National University of Singapore | Shanghai Jiao Tong University

To address the inconsistency and scene mismatch arising from the decoupled treatment of speaker extraction and diarization in complex overlapping speech, this paper proposes the first end-to-end jointly optimized framework that unifies frequency-domain speech separation with time-domain speaker activity annotation. The method integrates deep clustering, mask estimation, speaker activity detection, and waveform-level separation modules, supporting variable numbers of speakers and arbitrary overlap ratios. A bidirectional协同 mechanism enables mutual enhancement between extraction and diarization, breaking away from conventional cascaded pipelines. Evaluated on LibriMix, SparseLibriMix, and the real-world telephone conversation dataset CALLHOME, the approach achieves significant improvements over state-of-the-art methods on both tasks—marking the first demonstration of simultaneous gains in extraction quality (e.g., SI-SNRi) and diarization accuracy (e.g., DER).

Audio SeparationSpeaker DiarizationSpeaker Extraction

Latest Papers

What's happening recently
View more

This work proposes an end-to-end time-domain audio processing framework based on reservoir computing, addressing the limitations of traditional methods that rely on computationally intensive time–frequency transforms such as MFCCs and struggle to balance real-time performance, energy efficiency, and alignment with the human auditory system’s efficacy. By integrating biologically inspired auditory feature extraction with reservoir computing and replacing conventional frequency-domain transformations with lightweight convolutional operations, the proposed approach significantly reduces computational overhead while preserving discriminative feature representation. It eliminates the need for complex preprocessing and enables efficient, low-power real-time speech analysis, making it well-suited for embedded systems and voice-driven applications. This study thus establishes a highly energy-efficient and deployable paradigm for neuromorphic audio processing.

audio signal processingfeature extractionMFCC

This work proposes a knowledge-driven, self-supervised approach to audio segmentation and source separation that circumvents the reliance on large-scale manually annotated data. By integrating external prior knowledge—such as musical scores—into the audio processing pipeline for the first time, the method leverages hidden Markov models to achieve effective segmentation and separation of music and film audio without requiring labeled training data. Evaluated on synthetic datasets, the approach demonstrates strong performance, and in real-world film soundtrack tests, it significantly outperforms purely data-driven methods when incorporating sound-class priors. This represents a notable advance toward annotation-free audio analysis through principled integration of domain knowledge.

audio source separationcinematic audioknowledge-driven

This study addresses the lack of integrated analysis and visualization tools for high-dimensional acoustic features in bird vocalizations. The authors propose the first open-source, end-to-end framework that combines pYIN-based fundamental frequency estimation, MFCC extraction, and PCA dimensionality reduction to construct a unified timbral space, enabling high-quality audio reconstruction via the Griffin-Lim algorithm. Innovatively, the system integrates neural audio synthesis with a dual-view 3D interactive interface built on Three.js, supporting dynamic trajectory comparison and independent playback of original and synthesized vocalizations. Experimental results demonstrate a Mel-spectral correlation coefficient exceeding 0.92 in bird song reconstruction, confirming the framework’s high-fidelity preservation of perceptual acoustic structure.

3D visualizationacoustic analysisbioacoustics

This study investigates whether individual dimensions in the representations of self-supervised speech models (specifically WavLM) encode distinct speaker-related acoustic attributes, such as pitch, gender, intensity, noise level, and the second formant. By applying principal component analysis (PCA) to disentangle model features, the authors systematically identify independent dimensions that exhibit strong correlations with these acoustic properties, establishing for the first time a clear correspondence between specific latent dimensions and interpretable speaker characteristics. Further experiments demonstrate that manipulating these dominant dimensions enables effective control over the associated speaker attributes in speech synthesis, thereby confirming both their controllability and practical utility in downstream applications.

dimension analysisself-supervised speech featuresspeaker characteristics