CARES: A Controlled Synthetic Benchmark of Speaker Reactions to Sound

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of establishing reliable ground truth in audio scene description, where annotator subjectivity introduces significant ambiguity. To resolve this, we propose defining sound salience through speakers' audible reactions. Methodologically, we construct CARES, a synthetic benchmark corpus comprising ten thousand dual-speaker scenarios generated via controlled scene design and large language model-driven dialogue synthesis, thereby effectively eliminating annotation ambiguity. We evaluate six audio language models on this benchmark. Our experiments reveal that while existing models can recognize environmental sounds, they struggle to accurately capture speakers' behavioral responses to these sounds. This finding exposes a critical bottleneck in the field, highlighting the substantial gap between basic acoustic recognition and modeling human auditory interaction within complex audio scenes.
📝 Abstract
Automatic audio scene description turns a recording into a text account of a situation. One difficulty is deciding which elements of the audio should be kept, since a description cannot include them all. Annotators disagree about this, making a ground truth hard to obtain. In this work, we first define the ground truth, then generate the data. We focus on audio events and define sound salience with a simple rule: a sound is salient when a speaker audibly reacts to it. For scale and variety, a controlled set of scenarios fixes the ground truth, and a language model writes the dialogues. The resulting corpus, CARES, contains 10,000 two-speaker scenes. We then benchmark six audio-language models on three tasks: identifying the scene, tagging the sounds present, and classifying reactions. We show that these models hear the sounds but miss how the speakers react to them.
Problem

Research questions and friction points this paper is trying to address.

audio scene description
sound salience
ground truth
audio-language models
speaker reactions
Innovation

Methods, ideas, or system contributions that make the work stand out.

audio scene description
sound salience
synthetic benchmark
audio-language models
speaker reactions
🔎 Similar Papers
💼 Related Jobs
No related jobs found.