Logbook: Extremely Long-form Audio Event Understanding

πŸ“… 2026-10-05
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitation of existing audio benchmarks, which are confined to short clips and lack the capacity for hour-scale continuous audio event understanding. To this end, we introduce Logbook, the first benchmark supporting long-form audio understanding spanning hours to days, requiring models to perform gapless segmentation, label prediction, and caption generation on continuous recordings up to six days in length. We systematically evaluate 52 end-to-end and cascaded architectures, incorporating long-context processing and inference budget optimization strategies. Results demonstrate that end-to-end models generally outperform cascaded pipelines, though performance degrades as audio duration increases. Over-segmentation remains pervasive, and even the best-performing system falls substantially below human-level accuracy. These findings confirm the task’s feasibility while highlighting critical directions for future research.
πŸ“ Abstract
Audio benchmarks are built around short, pre-segmented clips, limiting model design to brief inputs or fixed vocabularies. To close this gap, we introduce Logbook, a benchmark for hour-scale audio understanding, with recordings ranging from ten minutes to six days. Given a continuous audio recording and an event label vocabulary, a system must predict a gap-free segmentation with an event label and a description per segment. We compare 52 systems, end-to-end and cascaded, and ablate fine-tuning, context length, and reasoning budget. We find the task tractable, though the best systems remain below the human reference. Also, over-segmentation is pervasive, and fine-tuning partially mitigates it. Finally, end-to-end are often better than cascaded systems, but degrades with longer context.
Problem

Research questions and friction points this paper is trying to address.

long-form audio understanding
audio event segmentation
audio benchmark
continuous audio recording
Innovation

Methods, ideas, or system contributions that make the work stand out.

Long-form Audio Understanding
Audio Event Segmentation
End-to-end Systems
Context Length
Benchmark
πŸ”Ž Similar Papers
2024-09-24IEEE International Conference on Acoustics, Speech, and Signal ProcessingCitations: 1