AVTrace: Diagnosing Audio-Visual Temporal Reasoning in Omni Models

📅 2026-09-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过引入AVTrace评估套件,诊断全模态模型在音频-视觉时间推理方面的能力,包括事件定位、同步判断等问题,并提出参数高效的时序后训练方法以改善模型性能。
📝 Abstract
Omni models can describe video content, but can they locate events in time, preserve event order, and judge audio-visual synchronization? We introduce AVTrace (Audio-Visual Temporal Reasoning Assessment and Capability Evaluation), a silver-standard diagnostic suite spanning onset and span grounding, synchronization, next-step prediction, cross-modal localization, chain parsing, and event-conditioned comprehension. It contains 34,114 training examples and category-balanced development and test splits of 3,500 and 7,000 examples. We evaluate five open omni models under their respective input configurations using reference-blind response normalization followed by deterministic scoring. All five off-the-shelf systems score below the test split's majority-label baseline of 0.556 on synchronization verification, and obtain low scores on chain parsing and event-conditioned grounding and comprehension. Development-set perturbations reveal task-dependent sensitivity in Qwen3-Omni-30B to modality removal and changes in visual input processing, without isolating their underlying causes. Parameter-efficient temporal post-training improves Gemma4-E4B-it on several benchmark metrics. On three external image benchmarks, task metrics change modestly, including some degradations, while teacher-forcing perplexity decreases. Together, these findings show that semantic reference-text overlap should not be treated as a proxy for temporal localization, and that AVTrace can identify task-specific weaknesses while providing a testbed for temporal post-training.
Problem

Research questions and friction points this paper is trying to address.

Audio-Visual Temporal Reasoning
Synchronization
Event Localization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Audio-Visual Temporal Reasoning
Omni Models
Synchronization Verification
Parameter-efficient Post-training
🔎 Similar Papers
No similar papers found.
L
Longyin Zhang
Institute of Advanced Intelligence and Computing (IAIC), A*STAR, Singapore
P
Parth Sakhare Mahendra
Institute of Advanced Intelligence and Computing (IAIC), A*STAR, Singapore
Chengwei Wei
Chengwei Wei
Research Scientist, Institute for Infocomm Research, A*STAR
Natural Language Processing
N
Ning Zhang
Institute of Advanced Intelligence and Computing (IAIC), A*STAR, Singapore
L
Lim Ming Chong
Institute of Advanced Intelligence and Computing (IAIC), A*STAR, Singapore
S
Sirui He
Institute of Advanced Intelligence and Computing (IAIC), A*STAR, Singapore
A
Ai Ti Aw
Institute of Advanced Intelligence and Computing (IAIC), A*STAR, Singapore