OmniCapBench: A Deep-Structured Evaluation Framework for Fine-Grained Audio-Visual Captioning

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inherent trade-off among coverage, localization precision, and scoring stability in existing audio-visual captioning evaluation benchmarks, which hinders precise diagnosis of multimodal large model capabilities. We propose a paradigm shift from free-text to atomized evaluation units, constructing a three-track structured framework encompassing entities, visual shots, and audio events. By integrating dense annotation decomposition, deterministic constraint verification, and localized LLM-based semantic comparison, our approach effectively decouples perceptual error types for fine-grained, reliable evaluation. Our analysis reveals that state-of-the-art models exhibit strong local perception yet weak long-range reasoning, with pronounced deficiencies in identity drift and cross-modal misalignment. These findings provide a precise roadmap for advancing omni-modal model development.
📝 Abstract
Multimodal large language models (MLLMs) are rapidly evolving toward continuous audio--visual reasoning, creating an urgent need for evaluations that expose their capability limits. Audio--visual captioning is an ideal diagnostic task, yet current benchmarks face a coupled trade-off: whole-caption scores provide coverage without localization, local probes provide localization without coverage, and unconstrained LLM judges introduce instability. We introduce OmniCapBench (Omni-Video Caption Benchmark), a benchmark that reframes audio--visual caption evaluation as a deep-structured diagnostic framework. OmniCapBench shifts the prediction target from free-form text to sets of atomic, verifiable evaluation units across three tracks: entity references, visual shots, and audio events, enabling reliable scoring with deterministic constraint checks and localized LLM-based semantic comparisons. With 786 densely annotated videos, OmniCapBench effectively distinguishes MLLM perception errors, including temporal grounding failures, identity drift, cross-modal misalignment, and hallucinated descriptions. Evaluating frontier MLLMs reveals strong local perception but weak long-horizon audio--visual reasoning, particularly in identity drift and cross-modal misalignment, providing a fine-grained roadmap for omnimodal development.
Problem

Research questions and friction points this paper is trying to address.

audio-visual captioning
multimodal large language models
evaluation benchmark
fine-grained assessment
cross-modal misalignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Audio-Visual Captioning
Evaluation Framework
Multimodal Large Language Models
Atomic Evaluation Units
Fine-Grained Diagnosis