audio event detection

Designs and builds algorithms and systems that detect, segment, and label events in audio streams, producing timestamps and confidence scores for occurrences such as speech, environmental sounds, or other acoustic events. This includes voice activity detection for speech/non-speech segmentation and broader sound-event-detection methods for multi-class identification and temporal localization of audio events.

audioeventdetection

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.16
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitations of conventional frame-level sound event detection models, which rely on post-processing and suffer from ambiguous temporal boundaries. To overcome these issues, the authors propose an end-to-end boundary-aware modeling approach that explicitly captures event start and end times through a Recurrent Event Detection (RED) layer and an Event Proposal Network (EPN). A tailored loss function is designed to enable boundary-sensitive optimization and inference. The method operates without requiring post-processing or hyperparameter tuning and achieves state-of-the-art performance across all classes on the AudioSet Strong annotation subset, significantly outperforming existing frame-level models and post-processing-based solutions.

Boundary-awareEvent BoundaryOnset and Offset

On Temporal Guidance and Iterative Refinement in Audio Source Separation

Jul 23, 2025
TM
Tobias Morocutti
🏛️ Institute of Computational Perception (CP-JKU) | LIT Artificial Intelligence Lab | Johannes Kepler University Linz

Traditional two-stage audio source separation approaches—first detecting sound events then separating sources—struggle in complex acoustic mixtures due to insufficient fine-grained temporal modeling. To address this, we propose a time-varying collaborative framework: (1) a fine-tuned pre-trained Transformer performs high-accuracy, frame-level sound event detection (SED) to generate dynamic temporal guidance signals; (2) an iterative refinement separation network explicitly models temporal dynamics by jointly incorporating label-conditioned constraints and recursive output feedback. Evaluated on DCASE 2025 Task 4, our method achieves second place, significantly improving both SED F1-score (+2.1%) and separation quality (SI-SNRi +1.8 dB) over strong baselines. Results demonstrate that time-aware collaborative modeling effectively bridges the gap between detection and separation, enabling more accurate and temporally coherent joint optimization.

Enhances synergy between event detection and separationImproves audio source separation with temporal guidanceUses iterative refinement to boost separation quality

High-cost manual annotation of sound event temporal boundaries severely limits the scalability of fully supervised learning; while existing weakly supervised methods predominantly rely on clip-level labels, partial label learning remains unexplored in audio analysis. This paper pioneers the introduction of partial label learning to audio event detection (AED), proposing a novel framework that leverages semantic priors from acoustic scene classification to automatically generate partial labels—i.e., annotating only a subset of positive time intervals for each sound event. We formulate a multi-task learning architecture jointly optimizing acoustic scene classification and AED. Furthermore, we design a self-distillation-based label refinement mechanism to synergistically train on both fully labeled and partially labeled data. Experiments demonstrate substantial reductions in annotation cost and consistent performance improvements across multiple benchmarks, validating the effectiveness and scalability of partial label learning in realistic audio scenarios.

Improving detection performance via semi-supervised training with partial labelsJointly analyzing acoustic scenes and sound events through multitask learningReducing annotation costs for sound event detection using partial labels

Existing methods for audio generation from silent videos lack fine-grained sound event labels—such as event type and onset time—that are temporally aligned with visual content, and often rely on post-processing steps that introduce errors. This work proposes the first unified framework that jointly models audio generation and sound event annotation by introducing an event-aware mechanism in the latent space, enabling end-to-end multi-task training to simultaneously produce audio and frame-level aligned event labels. Evaluated on the Greatest Hits dataset, the approach improves sound onset detection accuracy from 46.7% to 75.0% and boosts material classification accuracy across 17 categories from 40.6% to 61.0%, substantially enhancing the interpretability and practical utility of the generated audio.

audio event labelingmultimodal generationsilent video

Exploring Text-Queried Sound Event Detection with Audio Source Separation

Sep 20, 2024
HY
Han Yin
🏛️ Northwestern Polytechnical University | Alibaba Group | Fortemedia Singapore

To address performance degradation in sound event detection (SED) caused by overlapping acoustic events and background noise, this paper proposes a text-query-based SED framework (TQ-SED). First, we introduce AudioSep-DP—a language-driven, end-to-end differentiable audio separation model—by augmenting AudioSep with a dual-path RNN module to enhance dynamic audio modeling. Second, we design a multi-branch specialized detection head to independently identify events from each separated source. To the best of our knowledge, this is the first work to jointly model cross-modal language–audio alignment, a dual-path RNN–CNN hybrid separation architecture, and event detection. Evaluated on the DCASE 2024 Task 9 objective single-model track, TQ-SED achieves first place, improving F1 score by 7.22% over conventional SED methods. The code and pretrained models are publicly available.

Audio Event DetectionNoise ReductionSound Event Separation

Latest Papers

What's happening recently
View more

To address low accuracy and poor robustness in real-time multimodal anomaly detection caused by audio-video asynchrony and modality mismatch in industrial settings, this paper proposes a unified multimodal fusion framework. It introduces a bidirectional cross-modal attention mechanism for fine-grained alignment between video (YOLOv8/DETR + ByteTrack) and audio (AST/Wav2Vec2/HuBERT) streams, and integrates hybrid object detection with multi-strategy anomaly discrimination to enable synchronous streaming inference. Evaluated on general surveillance and industrial safety benchmarks, the system achieves >25 FPS on standard GPUs, with improvements of +4.2% mAP, +7.8% anomaly detection rate, and −12.3% false positive rate. Key contributions include: (i) the first end-to-end audio-visual synchronized anomaly detection system designed specifically for industrial deployment; and (ii) an extensible cross-modal attention architecture coupled with a lightweight audio integration scheme.

Creating industrial safety applications with real-time performance on standard hardwareDeveloping real-time multimodal anomaly detection using synchronized video and audio processingImproving accuracy and robustness through cross-modal attention and multi-model ensembles

FlexSED: Towards Open-Vocabulary Sound Event Detection

Sep 22, 2025
JH
Jiarui Hai
🏛️ Johns Hopkins University

Existing sound event detection (SED) methods are constrained by the closed-vocabulary assumption, limiting support for free-text queries and exhibiting poor zero-shot and few-shot generalization. Moreover, existing text-guided source separation techniques are ill-suited for SED tasks requiring fine-grained temporal localization and open-vocabulary retrieval. To address these limitations, we propose Open-SED—a novel framework for open-vocabulary SED. It integrates an audio self-supervised encoder (e.g., Data2Vec) with the CLAP text encoder, employs an adaptive cross-modal fusion decoder, and leverages large language models to generate high-quality, diverse event queries for enhanced supervision. Through continuous pretraining and collaborative annotation, Open-SED achieves substantial improvements on AudioSet-Strong: it outperforms conventional SED methods by +12.3% mAP in zero-shot and +9.7% mAP in 5-shot settings. The code and pretrained models are publicly released.

Detects sounds from free-text queriesEnables zero-shot and few-shot sound detectionImproves temporal localization for diverse sound vocabularies

This work addresses the limitations of existing sound event detection datasets, which often rely on synthetic or web-sourced audio and lack the diversity and annotation reliability of real-world home environments. To bridge this gap, the authors introduce a new benchmark dataset comprising 5,710 naturally recorded household audio clips spanning 15 common sound event classes. For the first time in this domain, a multi-annotator scheme coupled with a rigorous validation protocol is employed to ensure label quality, and rich metadata—including recording device, location, and environmental context—is provided. A Transformer-based baseline model, enhanced with strategies for annotation aggregation, post-processing, long-audio inference, and metadata fusion, achieves a macro-averaged PSDS1 score of 0.731 on the test set, establishing a high-quality benchmark and a robust evaluation framework for sound event detection in realistic domestic settings.

annotation variabilitydataset benchmarkdomestic environment

Region-Specific Audio Tagging for Spatial Sound

Sep 11, 2025
JZ
Jinzheng Zhao
🏛️ University of Surrey | Tencent AI Lab | University of Science and Technology Beijing | The Chinese University of Hong Kong

This work addresses the limitation of existing audio tagging methods, which cannot localize and tag sound events within specific spatial regions (e.g., designated azimuth angles or radial distances) in spatial audio. We formally introduce “region-specific audio tagging” as a novel task. Methodologically, we propose a multimodal feature representation that jointly encodes spectral, spatial directional, and positional information; extend pre-trained models—PANNs and AST—into spatially aware architectures; and incorporate directional feature enhancement to improve omnidirectional tagging capability. Experiments on both simulated and real microphone array datasets demonstrate substantial improvements in region-specific sound source identification accuracy. Results validate both the well-posedness of the proposed task and the effectiveness of our technical approach. This work establishes a new paradigm for spatial audio understanding and provides a scalable, foundational framework for future research in spatially grounded audio analysis.

Combining spectral, spatial, and position featuresExtending audio tagging systems for directional audioLabeling sound events in specific spatial regions

This work addresses the limitations of existing approaches in long-form audio activity recognition, which suffer from inconsistent cross-level modeling and reliance on multi-level supervisory labels. The authors reformulate the task as a hierarchical parsing problem grounded in event-level evidence, introducing a hierarchical activity grammar to enforce compositional structure and temporal ordering constraints. They further propose a grammar-guided dynamic programming decoding mechanism that, using only event-level detection posteriors, enables end-to-end generation of temporally coherent and semantically interpretable activity–sub-activity–event parse trees—without requiring supervision from high-level activity or sub-activity annotations. Evaluated on the MultiAct dataset, the method achieves substantial improvements in temporal consistency, as measured by Edit Score, while enabling explainable inference of hierarchical activity structures.

activity recognitioncross-level inconsistencyhierarchical parsing

Hot Scholars

HY

Hung-yi Lee

National Taiwan University
deep learningspoken language understandingspeech processing
SW

Shinji Watanabe

Carnegie Mellon University
Speech recognitionSpeech processingSpeech enhancementSpeech translation
CY

Chun-Yi Kuan

National Taiwan University
Speech ProcessingDeep LearningSpoken Language Understanding
ZZ

Zhou Zhao

Zhejiang University
Machine LearningData MiningMultimedia Computing