Revisiting Frame-Wise Saliency for Audio Moment Retrieval

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited prediction performance of decoders in audio moment retrieval by proposing a parameter-free segmentation method. Instead of relying on conventional decoder outputs, this approach directly generates and ranks candidate moments using frame-level saliency sequences. Building upon a DETR-based architecture and segmentation rules inspired by Sound Event Detection (SED), this work is the first to demonstrate that saliency sequences from auxiliary outputs can serve as an effective prediction source, revealing their distinct roles in training supervision versus inference prediction. Experimental results show that the proposed method yields significant improvements in the R1@0.7 metric on datasets such as CASTELLA, with particularly notable gains in short-moment retrieval.
📝 Abstract
This paper revisits frame-wise saliency for audio moment retrieval (AMR). We show that the frame-wise saliency sequence, conventionally used only as an auxiliary output in DETR-based AMR models, can itself serve as an effective source of moment predictions. We convert the saliency sequence into ranked moments using a simple SED-inspired segmentation rule with no learned parameters, enabling moment retrieval directly from frame-wise temporal information. On the CASTELLA dataset, saliency-based prediction consistently outperforms decoder-based prediction from the same model across all 18 runs of QD-DETR and CG-DETR. For QD-DETR, simply replacing the inference output improves R1@0.7 from 21.0 to 36.1. The advantage remains 7-15 points when the two outputs are evaluated at their independently selected best epochs. The same tendency extends to TaskWeave and UVCOM, whereas TR-DETR shows the opposite behavior, suggesting that how saliency construction may matter. The performance gap is especially pronounced for short moments: for queries whose annotated moments average at most 2 s, R1@0.7 improves from 6.3 to 27.8 with QD-DETR. Decoder supervision nevertheless benefits saliency-based prediction, indicating that its role during training differs from the utility of its inference output.
Problem

Research questions and friction points this paper is trying to address.

Audio Moment Retrieval
Frame-Wise Saliency
DETR-based models
Short moment retrieval
Innovation

Methods, ideas, or system contributions that make the work stand out.

Audio Moment Retrieval
Frame-wise Saliency
DETR-based Models
Parameter-free Segmentation
Short Moment Retrieval
🔎 Similar Papers
2024-09-24IEEE International Conference on Acoustics, Speech, and Signal ProcessingCitations: 1
2024-07-18IEEE Workshop/Winter Conference on Applications of Computer VisionCitations: 0