π€ AI Summary
This study addresses the limitation of existing audio benchmarks, which are confined to short clips and lack the capacity for hour-scale continuous audio event understanding. To this end, we introduce Logbook, the first benchmark supporting long-form audio understanding spanning hours to days, requiring models to perform gapless segmentation, label prediction, and caption generation on continuous recordings up to six days in length. We systematically evaluate 52 end-to-end and cascaded architectures, incorporating long-context processing and inference budget optimization strategies. Results demonstrate that end-to-end models generally outperform cascaded pipelines, though performance degrades as audio duration increases. Over-segmentation remains pervasive, and even the best-performing system falls substantially below human-level accuracy. These findings confirm the taskβs feasibility while highlighting critical directions for future research.
π Abstract
Audio benchmarks are built around short, pre-segmented clips, limiting model design to brief inputs or fixed vocabularies. To close this gap, we introduce Logbook, a benchmark for hour-scale audio understanding, with recordings ranging from ten minutes to six days. Given a continuous audio recording and an event label vocabulary, a system must predict a gap-free segmentation with an event label and a description per segment. We compare 52 systems, end-to-end and cascaded, and ablate fine-tuning, context length, and reasoning budget. We find the task tractable, though the best systems remain below the human reference. Also, over-segmentation is pervasive, and fine-tuning partially mitigates it. Finally, end-to-end are often better than cascaded systems, but degrades with longer context.