Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos

📅 2026-07-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing audio-visual large models struggle to effectively comprehend long-duration, complex real-world content. To address this limitation, this work proposes an open audio-visual large language model that jointly models cross-modal, temporal, and compositional semantics in long videos. The approach leverages a newly curated Audio-Visual-Skills dataset, employs a three-stage progressive curriculum training strategy, and introduces a temporally aligned audio-visual interleaved chain-of-thought reasoning framework. This methodology substantially enhances both performance and interpretability in long-form video understanding, outperforming existing open-source models by significant margins across more than fifteen multimodal benchmarks. Notably, it even surpasses certain larger closed-source systems on specific tasks, demonstrating exceptional generalization capability and practical potential.
📝 Abstract
We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, and long-form videos. Unlike prior AV-LLMs that primarily focus on short clips, AV-Flamingo is designed for understanding and reasoning over long and complex real-world (audio-visual) videos. To support this, we make three key contributions: (i) Audio-Visual-Skills, a large-scale collection of real-world videos with ~7M caption and question-answer training instances designed to emphasize temporal, compositional, and cross-modal audio-visual reasoning; (ii) a novel three-stage curriculum that progressively trains the model from short-range perception to long-horizon multi-event reasoning; and (iii) Temporal Audio-Visual Interleaved Chain-of-Thought, a reasoning framework that explicitly grounds intermediate reasoning steps to timestamps in long audio-visual streams, improving temporal alignment and interpretability. Extensive experiments across 15+ audio-visual, omni-modal, audio, and vision benchmarks show that AV-Flamingo outperforms similarly sized open models by clear margins and remains highly competitive with, and in some cases surpasses, much larger open-weight and closed models, particularly on long and complex real-world audio-visual understanding and reasoning tasks. Beyond benchmark performance, AV-Flamingo exhibits strong real-world utility and transfers well to unseen tasks, highlighting its robustness and generalization ability.
Problem

Research questions and friction points this paper is trying to address.

audio-visual understanding
long-form videos
complex videos
temporal reasoning
cross-modal reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

audio-visual large language model
long-form video understanding
temporal reasoning
cross-modal alignment
curriculum learning
🔎 Similar Papers
2024-07-18IEEE Workshop/Winter Conference on Applications of Computer VisionCitations: 0
Sreyan Ghosh
Sreyan Ghosh
Ph.D. in CS at University of Maryland, College Park
AIMachine LearningNLPSpeech Recognition
Arushi Goel
Arushi Goel
Research Scientist, NVIDIA
Computer VisionMachine LearningVision and Language
K
Kaousheik Jayakumar
University of Maryland, USA
L
Lasha Koroshinadze
University of Maryland, USA
Nishit Anand
Nishit Anand
MS CS at University of Maryland, College Park
Machine LearningComputer VisionNatural Language ProcessingSpeech Recognition
Siddharth Gururani
Siddharth Gururani
NVIDIA Research
Artificial IntelligenceMusic Information RetrievalMachine LearningDeep LearningText to Speech
Hanrong Ye
Hanrong Ye
NVIDIA Research
multi-task multi-modal models
P
Pritam Biswas
NVIDIA, USA
Y
Yuanhang Su
NVIDIA, USA
Ehsan Hosseini-Asl
Ehsan Hosseini-Asl
Senior Deep Learning Scientist at Nvidia
Deep LearningNLPGenerative AIComputer VisionSpeech
Sang-gil Lee
Sang-gil Lee
NVIDIA
Deep Generative ModelAudio SynthesisLanguage Model
Zhifeng Kong
Zhifeng Kong
Senior Research Scientist, NVIDIA
Deep Generative ModelsDiffusion ModelsAudio Foundation ModelsAudio LMTrustworthy ML
Jaehyeon Kim
Jaehyeon Kim
NVIDIA
Machine Learning
S
Sungwon Kim
NVIDIA, USA
S Sakshi
S Sakshi
Ph.D. in CS at University of Maryland, College Park
Machine LearningNatural Language ProcessingAudio Processing
Ramani Duraiswami
Ramani Duraiswami
Computer Science and UMIACS, University of Maryland
Scientific ComputingSpatial AudioMachine LearningComputational Electromagnetics
Dinesh Manocha
Dinesh Manocha
Distinguished University Professor, University of Maryland at College Park
computer graphicsgeometric modelingmotion planningvirtual realityrobotics
Andrew Tao
Andrew Tao
Nvidia
Computer VisionMachine Learning
Mohammad Shoeybi
Mohammad Shoeybi
Senior Director of Applied Research at NVIDIA
Large Language ModelsNLPMulti-Modal ModelsGenerative AI
Bryan Catanzaro
Bryan Catanzaro
NVIDIA
Parallel ComputingMachine Learning
Ming-Yu Liu
Ming-Yu Liu
Vice President of Research at NVIDIA, IEEE Fellow
Computer VisionMachine Learning
Wei Ping
Wei Ping
Distinguished Research Scientist, NVIDIA
machine learninglarge language modelsspeech synthesisreinforcement learning