CultureVidBench: Benchmarking Cultural Understanding in Text-to-Video Generation

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the lack of fine-grained evaluation of text-to-video (T2V) generation models in diverse cultural contexts by introducing CultureVidBench—the first cross-cultural benchmark for video generation. Spanning 12 countries, 6 continents, 8 cultural regions, and 14 cultural dimensions, it comprises 1,000 dynamic multimodal prompts focused on the faithful representation of material culture, social practices, and ceremonial rituals. We develop a fine-grained cultural understanding evaluation framework that integrates human user studies with automatic assessment via multimodal large language models (MLLMs), evaluating performance across four dimensions: cultural fidelity, multimodal rendering, semantic consistency, and perceptual quality. Experiments reveal that while mainstream T2V models achieve strong visual quality, they frequently exhibit cultural distortions—particularly in underrepresented regions, complex rituals, and nuanced multimodal cultural details.
📝 Abstract
Text-to-video (T2V) generation models have advanced rapidly, yet their ability to represent diverse cultural contexts remains underexplored. Existing benchmarks mainly focus on perceptual quality, physical plausibility, and text-video alignment, but do not directly assess whether generated videos capture culturally specific objects, actions, rituals, visible text, or audio cues. We introduce CultureVidBench, a comprehensive benchmark for evaluating cultural understanding in T2V generation. CultureVidBench contains 1,000 curated prompts covering 12 countries, 6 continents, 8 cultural regions, and 14 cultural aspects organized into three categories: material culture, social practice & performance, and ritual & ceremony. Designed specifically for video generation, CultureVidBench emphasizes dynamic and multimodal cultural representation, including social interactions, ritual procedure, and culturally appropriate visible text and audio. We evaluate seven representative T2V models through human user studies and MLLM-based automatic assessment across cultural faithfulness, multimodal cultural rendering, semantic adherence, and perceptual quality. Results show that although current models achieve strong semantic adherence and visual quality, they often fail to faithfully capture fine-grained cultural details, particularly for underrepresented regions, rituals, and multimodal cultural cues.
Problem

Research questions and friction points this paper is trying to address.

cultural understanding
text-to-video generation
benchmark
multimodal cultural representation
cultural faithfulness
Innovation

Methods, ideas, or system contributions that make the work stand out.

cultural understanding
text-to-video generation
multimodal benchmark
CultureVidBench
MLLM-based evaluation