How Does Audio Influence Visual Attention in Omnidirectional Videos? Database and Model

📅 2024-08-10
🏛️ arXiv.org
📈 Citations: 1
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates how audio modulates visual attention in omnidirectional video (ODV) to enhance VR/AR user experience. To this end, we introduce AVS-ODV—the first large-scale audio-visual joint eye-tracking dataset for ODV—comprising 162 panoramic videos, 60 participants, and three audio conditions. It enables the first empirical characterization of Ambisonics audio’s directional modulation effect on ODV attention distributions. We propose OmniAVS, a hierarchical multimodal alignment-fusion model integrating spherical projection modeling, cross-modal embedding space alignment, and a U-Net-based decoder, overcoming limitations of unimodal saliency prediction. Experiments demonstrate that OmniAVS achieves state-of-the-art performance on audio-visual saliency prediction for ODV. Both the AVS-ODV dataset and OmniAVS code are publicly released, establishing a new benchmark and open resource for multimodal attention modeling.

Technology Category

Computer Vision: Multi-modal VisionMachine Learning: Multimodal LearningIntelligent Robots: Multimodal Perception & Sensor Fusion

Application Category

User Modeling, Personalization and Recommendation: On-Device user modeling, personalization, and recommendationSearch and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGSecurity and Privacy: Data transparency and provenance
📝 Abstract
Understanding and predicting viewer attention in omnidirectional videos (ODVs) is crucial for enhancing user engagement in virtual and augmented reality applications. Although both audio and visual modalities are essential for saliency prediction in ODVs, the joint exploitation of these two modalities has been limited, primarily due to the absence of large-scale audio-visual saliency databases and comprehensive analyses. This paper comprehensively investigates audio-visual attention in ODVs from both subjective and objective perspectives. Specifically, we first introduce a new audio-visual saliency database for omnidirectional videos, termed AVS-ODV database, containing 162 ODVs and corresponding eye movement data collected from 60 subjects under three audio modes including mute, mono, and ambisonics. Based on the constructed AVS-ODV database, we perform an in-depth analysis of how audio influences visual attention in ODVs. To advance the research on audio-visual saliency prediction for ODVs, we further establish a new benchmark based on the AVS-ODV database by testing numerous state-of-the-art saliency models, including visual-only models and audio-visual models. In addition, given the limitations of current models, we propose an innovative omnidirectional audio-visual saliency prediction network (OmniAVS), which is built based on the U-Net architecture, and hierarchically fuses audio and visual features from the multimodal aligned embedding space. Extensive experimental results demonstrate that the proposed OmniAVS model outperforms other state-of-the-art models on both ODV AVS prediction and traditional AVS predcition tasks. The AVS-ODV database and OmniAVS model will be released to facilitate future research.
Problem

Research questions and friction points this paper is trying to address.

Investigates audio-visual attention in omnidirectional videos (ODVs)
Introduces new AVS-ODV database with 162 ODVs and eye-tracking data)
Proposes OmniAVS model for superior audio-visual saliency prediction)
Innovation

Methods, ideas, or system contributions that make the work stand out.

Introduces AVS-ODV database with 162 omnidirectional videos
Proposes OmniAVS network for audio-visual saliency prediction
Hierarchically fuses audio and visual features effectively
🔎 Similar Papers
No similar papers found.