Misrecognition or Abstraction? Rethinking Outputs of Sound Event Recognition

📅 2026-09-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对声音事件识别在不确定性下的输出问题,提出了一种结合声音事件类别、置信度和拟声描述的新方法,以提高信息量和沟通性。
📝 Abstract
Conventional general sound recognition systems typically output deterministic sound event labels, implicitly assuming that the target sound class can be correctly identified from the input audio. However, in real listening situations, the sound event class is not always clearly identifiable. Human listeners may nevertheless understand their surroundings from an ambiguous sound without identifying its exact sound event class. This motivates a discussion of how the outputs of sound recognition systems should be redesigned under such uncertainty. As a basis for this discussion, this paper proposes an output representation for sound event recognition that combines a sound event class, its confidence score, and an onomatopoeic description of the sound. The proposed representation preserves conventional class-based recognition while providing an additional onomatopoeic description of acoustic characteristics that can remain informative even when the class prediction is uncertain. Experiments using ESC-50 and ESC-50-Onomatopoeia show that the proposed method achieves sound recognition performance comparable to that of a conventional recognition-only system. In addition, an LLM-as-a-judge evaluation and subjective listening experiments indicate that the proposed output is preferred over conventional deterministic outputs based on the sound event label, particularly when used to support understanding of the surrounding environment. These results suggest that such output representations can make sound event recognition more informative and communicative under uncertainty.
Problem

Research questions and friction points this paper is trying to address.

Sound Event Recognition
Uncertainty
Output Representation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sound Event Recognition
Confidence Score
Onomatopoeic Description
🔎 Similar Papers
No similar papers found.
N
Naoya Tomida
Kyoto University, Japan
Y
Yuki Okamoto
The University of Tokyo, Japan
Keisuke Imoto
Keisuke Imoto
Kyoto University
Acoustic Signal ProcessingSound Event Detection