SsgCaps: A controlled dataset for the evaluation of sound scene generation algorithms

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文介绍了SsgCaps数据集,用于评估声音场景生成算法。通过与原版数据集对比分析,证明其开放版本适合进一步基准测试。
📝 Abstract
Sound Scene Generation is about the automatic synthesis of artificial sound scenes. We introduce SsgCaps, a publicly available dataset of human-engineered sound scenes wherein each scene matches a precisely structured prompt that guides the sampling process. The corresponding prompts are sampled from a predefined action-based typology that allows extensive sampling while retaining plausibility. SsgCaps is a sound scene dataset derived from the unpublished reference dataset for Task 7 of the 2024 DCASE Challenge edition, which contained private-and public-domain audio samples. In contrast, SsgCaps contains only public-domain audio samples, allowing us to open this dataset to the community. To make this dataset useful to the community, we first elaborate on the rationale for the prompt and dataset structure. We then perform a comparative quantitative analysis of the 2 versions of the dataset. To do so, we compare both versions to the audio synthesized by the SSG algorithms submitted to the challenge using Fr{é}chet Audio Distance (FAD) and Kernel Audio Distance (KAD) as well as perceptual ratings. This analysis shows only small differences, which enables us to recommend the open version for further benchmarking of SSG algorithms.
Problem

Research questions and friction points this paper is trying to address.

Sound Scene Generation
Dataset
Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

SsgCaps
sound scene generation
public-domain audio samples
Fréchet Audio Distance (FAD)
Kernel Audio Distance (KAD)
🔎 Similar Papers
No similar papers found.