LEAP: Learned Block-wise Evidence Retrieval for Long Audio-Video Perception

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the dilemma in hour-long audio-visual question answering (AVQA) where dense encoding exhausts context windows while uniform compression dilutes fine-grained evidence. We propose LEAP, a framework that decouples evidence localization from reasoning. It employs lightweight chunked retrieval to identify high-confidence windows and performs bounded reasoning exclusively on critical segments, rendering context overhead independent of total video duration and enabling streaming causal queries. Furthermore, candidate window score pooling and two-stage LoRA fine-tuning preserve native audio-visual details within multimodal large language models. Experiments demonstrate that LEAP outperforms baselines by 4.5–16.8% across multiple AVQA benchmarks, and maintains gains of 3.1–13.0% over published results when transferred across different backbones.
📝 Abstract
Hour-scale audio-visual question answering is constrained by a context dilemma: dense whole-recording encoding rapidly exhausts context limits, whereas uniform temporal compression severely dilutes fine-grained acoustic and visual evidence. We introduce LEAP, a framework where the model retrieves its own evidence without placing the whole recording in one context. LEAP divides a recording into fixed-duration blocks, applying a lightweight localization pass to each block to score short candidate windows. The highest-ranked windows are pooled and re-encoded in a single bounded answer pass. Consequently, the answer input and peak context remain independent of the recording duration. By decoupling evidence localization from reasoning, our framework can localize candidate temporal windows over pre-computed transcripts without decoding media frames, while preserving fine-grained visual and non-speech evidence by routing the final answering pass over raw audio-visual streams. LEAP trains both stages: a localization LoRA improves the selected windows, and an answer LoRA improves the answers read from the same windows. The block grid natively supports causal queries, enabling LEAP to support streaming inference without streaming-specific training. Across several AVQA benchmarks, LEAP improves over the Qwen3-Omni-30B-A3B baseline by 4.5-16.8%, and transfers to a second omni-modal backbone, MiniCPM-o 4.5, surpassing its published results by 3.1-13.0%.
Problem

Research questions and friction points this paper is trying to address.

audio-visual question answering
long audio-video perception
context dilemma
evidence retrieval
hour-scale AVQA
Innovation

Methods, ideas, or system contributions that make the work stand out.

block-wise evidence retrieval
audio-visual question answering
decoupled localization-reasoning
dual LoRA training
streaming inference
💼 Related Jobs
No related jobs found.