curate audio-visual dataset

Designs and builds curated multimodal datasets that pair synchronized audio and visual signals with structured annotations and metadata (e.g., transcriptions, phoneme/G2P alignments, on-screen sound labels, and editing-triplet groupings). This competence covers specifying collection and synthetic-generation pipelines, producing category-balanced and human-centric examples, constructing on-screen-sound and other AV subsets, and implementing multi-pipeline and agent-in-the-loop quality-control, validation, and evaluation processes.

curateaudio-visualdataset

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.15
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Effectively obtaining acoustic, visual and textual data from videos

Sep 06, 2025
JE
Jorge E. León
🏛️ Adolfo Ibáñez University (UAI) | Diego Portales University (UDP)

A critical shortage exists of high-quality, large-scale, semantically aligned acoustic–visual–textual multimodal datasets. Method: This paper proposes an end-to-end framework for constructing video-based multimodal data, integrating three key components: (i) video content filtering, (ii) cross-modal synchronization triplet extraction (audio–frame–subtitle), and (iii) fine-grained description synthesis leveraging image-to-text generation models—ensuring temporal and semantic alignment across all three modalities. Contribution/Results: The resulting publicly released dataset spans diverse real-world scenarios and substantially advances performance on cross-modal retrieval and joint embedding learning tasks, achieving state-of-the-art results across multiple benchmarks. By providing a scalable, high-fidelity resource, this work establishes a new foundation for training and evaluating foundational multimodal models.

Creating high-quality audio-image-text datasetsEnsuring semantic connections between modalitiesExtracting multimodal data from videos

Audio-Language Datasets of Scenes and Events: A Survey

Jul 09, 2024
GW
Gijs Wijngaard
🏛️ Maastricht University

This study systematically evaluates 69 audio-language datasets available as of September 2024, revealing pervasive issues including acoustic class imbalance, multi-source duplication, linguistic homogeneity (dominant English bias), restricted accessibility, and latent societal biases. Methodologically, we innovatively integrate PCA-based cross-dataset embedding variance analysis, CLAP-guided detection of modality leakage, joint acoustic–textual distribution modeling, and open governance practices to quantitatively identify systemic biases—particularly in widely used sources such as YouTube and Freesound. As a key contribution, we release an open resource library comprising over two million samples and propose a comprehensive Audio-Language Modeling (ALM) data curation roadmap that explicitly balances diversity, robustness, and fairness. This work establishes an empirically grounded, reproducible methodology for dataset development, directly supporting improved generalization capabilities of multimodal models.

Audio-lingual Model TrainingData Bias and LimitationsDataset Analysis

This work addresses the limitations of existing video generation models, which often neglect audio and rely on cascaded pipelines, leading to high computational costs, error propagation, and audio-visual desynchronization. To overcome these challenges, we propose MOVA—the first open-source, end-to-end image-and-text-to-video-and-audio (IT2VA) generation model. Built upon a 32B-parameter Mixture-of-Experts architecture (with 18B activated per forward pass), MOVA supports LoRA-based fine-tuning, efficient inference, and prompt enhancement. It simultaneously generates semantically aligned, high-fidelity video and audio, including lip-synced speech, contextually appropriate sound effects, and background music. By releasing the model weights and a complete toolchain, this work aims to advance research in joint audio-visual synthesis.

multimodal modelingopen-sourcescalability

Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Dec 19, 2024
HK
Ho Kei Cheng
🏛️ University of Illinois | Sony AI | Sony Group Corporation

To address poor audio quality, weak semantic alignment, and audio-visual desynchronization in video-to-audio generation, this paper proposes MMAudio, a multimodal joint-training framework. MMAudio is the first to unify video-audio and text-audio dual-path generation within a single architecture. It introduces a frame-level conditional synchronization module to achieve fine-grained alignment between video features and the audio latent space, and employs flow matching as the end-to-end optimization objective. The method supports either video-only or video-plus-text conditional inputs. On public benchmarks, MMAudio achieves state-of-the-art performance: significantly improved audio fidelity, enhanced semantic alignment, and reduced audio-visual synchronization error. At inference, it generates 8-second audio clips in 1.23 seconds, with a compact model size of only 157 million parameters.

Achieve state-of-the-art video-to-audio generation efficientlyImprove audio-visual synchronization via frame-level alignmentSynthesize high-quality audio from video and text

Read, Watch and Scream! Sound Generation from Text and Video

Jul 08, 2024
YJ
Yujin Jeong
🏛️ NAVER AI Lab

Existing video-to-audio generation methods suffer from limitations in sound source localization, semantic controllability, and audio fidelity consistency. To address these, we propose the first high-fidelity, video–text joint-driven audio synthesis framework. Our method employs a decoupled conditional injection mechanism: video inputs model acoustic structure (e.g., energy envelope), while text inputs encode semantic content—enabling independent user control over source intensity, ambient atmosphere, and primary sound semantics. Built upon a lightweight diffusion-based multimodal fusion architecture, it integrates a pretrained text-to-audio model with a video-derived energy estimation module, trained efficiently on audio–video–text triplets. Experiments demonstrate state-of-the-art performance: +1.2 MOS improvement in audio quality, −38% FID reduction in controllability, and 2.1× faster convergence. Code and real-time demo are publicly available.

Controllable High-Quality Audio SynthesisTargeted Sound GenerationVideo-to-Audio Conversion

Latest Papers

What's happening recently
View more

Existing Foley sound datasets generally suffer from insufficient quality and coarse annotations, hindering data-driven research in classification, retrieval, and synthesis. To address this gap, this work introduces and publicly releases FoleySet—a large-scale Foley dataset comprising 10,000 audio clips meticulously recorded following professional Foley practices. The dataset features a two-tier manual semantic annotation scheme that precisely aligns synchronous sound effects with on-screen human actions, such as footsteps, clothing rustles, and prop manipulations. FoleySet is the first to offer multi-level annotations, standardized formatting, and a permissive Creative Commons license, thereby filling a critical resource void in the field. It provides strong support for Foley-related audio tasks and advances research toward automated audiovisual content production.

annotated datasetaudiovisual post-productiondata scarcity

This work addresses the limitations of existing audio-to-image generation methods, which struggle to effectively fine-tune advanced text-to-image models due to the scarcity of high-quality, cross-modally aligned datasets. To overcome this challenge, the authors introduce A2I-Set, the first large-scale trimodal-aligned dataset comprising 323,000 triples of audio clips, images, and fine-grained textual descriptions, alongside a dedicated generative framework named AudioCanvas. AudioCanvas is fine-tuned on A2I-Set to enable audio-driven image synthesis. The study further proposes a human-supervised, mixed-source test set to rigorously evaluate cross-modal consistency. Experimental results demonstrate that AudioCanvas significantly outperforms current approaches in both visual fidelity and semantic alignment between audio and generated images, thereby validating the effectiveness of the proposed dataset and model architecture.

audio-to-image generationcross-modal alignmentexpressive synthesis

This study addresses the critical limitation in sound effects research caused by incompatible labeling schemes and metadata structures across existing datasets, which hinders data integration and cross-study comparison in both classification and generation tasks. To overcome this, the authors propose the first universal category system (UCS)-based relabeling framework tailored for academic use. The framework employs a rule-driven, multi-stage pipeline to unify heterogeneous labels across datasets, incorporating mechanisms for conflict resolution, hierarchical categorization, and cross-source alignment. By applying the industry-standard UCS to academic sound data for the first time, the work constructs and publicly releases EnvSound-UCS—a harmonized dataset integrating 58,057 samples from AudioSet, FSD50K, and ESC-50. This resource achieves high automated conversion rates while substantially improving label consistency, effectively mitigating sound data fragmentation.

data silosdataset unificationmetadata heterogeneity

This work addresses the limitations of existing audio-visual synchronization evaluation methods, which struggle to disentangle temporal alignment from semantic consistency and suffer from coupling biases in data construction. We propose the first structured and scalable benchmark framework that enables independent assessment of temporal synchronization and semantic correspondence. Through a hybrid pipeline combining automated filtering and human verification, we construct a large-scale dataset comprising 3,269 videos and 38,390 samples across three audio categories—speech, music, and environmental sounds—and ten diverse scenarios. The dataset ensures authentic on-screen sound sources and supports both multimodal alignment analysis and downstream task evaluation. Using this benchmark, we systematically evaluate five representative models. Both code and data are publicly released.

audio-visual synchronizationevaluation benchmarkmultimodal understanding

This work addresses the challenges hindering systematic progress in audio-visual intelligence—namely, task heterogeneity, inconsistent taxonomies, and fragmented evaluation protocols—by proposing the first unified task taxonomy encompassing understanding, generation, and interaction. It introduces a cohesive methodological framework that integrates key techniques including modality tokenization, cross-modal fusion, autoregressive and diffusion-based generation, large-scale pretraining, instruction alignment, and preference optimization. Furthermore, the study establishes a structured repository of datasets, benchmarks, and comparative evaluation metrics, while systematically identifying critical open challenges such as temporal synchronization, spatial reasoning, controllability in generation, and safety considerations, thereby offering a unified reference and clear technical roadmap for future research in the field.

Audio-Visual Intelligenceevaluation heterogeneityfoundation models

Hot Scholars

MB

Martijn Bartelds

Postdoctoral Scholar, Stanford University
Computational linguisticsLanguage variationSpeech technology
TF

Teng Fei

School of Resources and Environmental Science, Wuhan University
Remote SensingGISSocial SensingPlanning
RY

Ran Yi

Associate Professor, Shanghai Jiao Tong University
Computer VisionComputer Graphics
OA

Orevaoghene Ahia

University of Washington
Natural Language ProcessingComputational Linguistics
SW

Shangda Wu

Tencent
Symbolic Music GenerationMusic Information RetrievalMultimodal Learning