InteracVid: Building a Real Interactive Audio-Visual Response Dataset from Live-Chat Videos

📅 2026-08-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing audio-visual generation models struggle to produce natural and causally coherent multimodal responses due to the absence of real-world user interaction supervision data. This work proposes and constructs InteracVid, the first large-scale open-source interactive audio-visual dataset, by extracting context–stimulus–response triplets from live-stream videos across diverse scenarios including dialogue, physical manipulation, and demonstrations. Leveraging a metadata-aware mining pipeline, the authors automatically curate 454K high-quality triplets from over 59K videos amidst noisy live-stream content, explicitly encoding causal structure, temporal completeness, and naturalness. Experiments demonstrate that fine-tuning on this dataset substantially improves interactive planning and response generation quality, with consistent gains observed across both human evaluation and automatic metrics on a benchmark of 100 real-world chat queries.
📝 Abstract
Large language models have made text the default medium for human--AI interaction, buttext alone cannot express the full range of responses required by multimodal assistants,avatars, and embodied agents. While recent audio-video generative models can synthesizehigh-fidelity synchronized content, existing supervision is largely \emph{descriptive}:models are trained to render captions rather than to produce audio-visual responsescaused by external user interactions. We introduce \textbf{InteracVid}, \emph{the firstopen-source large-scale dataset that addresses this missing supervision}, so that everysample couples a preceding audio-visual context and an external stimulus with the realinteractive response that follows. We design a metadata-aware pipeline that extractsinteractive clips from long, noisy livestreams, yielding over \textbf{454K}context-query-response triplets from more than \textbf{59K} livestream videos andspanning conversation-centered, object-centric, procedural, embodied, and screen-basedscenarios. A ten-rater human study confirms that the extracted interactions are causal,natural, and temporally complete for both genuine and reconstructed queries. On aheld-out benchmark of \textbf{100} genuine live-chat queries, fine-tuning on InteracVidimproves both interaction planning and audio-video response generation, and anindependent human evaluation reproduces the system ranking and the conclusions obtainedwith our automatic judge. These results highlight interaction-structured data as acritical foundation for interactive multimodal generation.
Problem

Research questions and friction points this paper is trying to address.

interactive multimodal generation
audio-visual response
causal supervision
live-chat videos
multimodal interaction
Innovation

Methods, ideas, or system contributions that make the work stand out.

interactive multimodal generation
audio-visual response dataset
causal interaction supervision
live-chat video mining
context-query-response triplets
🔎 Similar Papers
2024-06-09Annual Meeting of the Association for Computational LinguisticsCitations: 13