🤖 AI Summary
This study addresses the limitations of existing video safety review systems, which predominantly adopt an adult-centric perspective and struggle to detect fine-grained, implicit, and context-dependent risks in AI-generated content (AIGC) targeting children. To bridge this gap, the authors introduce CAVSR—the first child-oriented benchmark for AIGC video risk assessment—and propose QVRS-E, a knowledge-enhanced multi-agent collaborative reasoning framework. QVRS-E integrates domain expertise with vision-language models to enable targeted evidence collection and fact-driven risk judgment. Moving beyond conventional generic violation detection paradigms, this approach significantly improves the accuracy of identifying child-related risks and demonstrates strong effectiveness and interpretability on a dataset of 605 real-world videos.
📝 Abstract
The rapid growth of Artificial Intelligence-generated content (AIGC) is reshaping video production and circulation, exposing children to an increasing volume of AIGC videos. Unlike traditionally produced videos, AIGC videos often exhibit greater uncertainty in visual details, narrative coherence, and content expression, which may introduce developmentally inappropriate risks for children. However, existing video safety research is largely designed for general violation detection from an adult perspective and remains insufficient for identifying the fine-grained, implicit, and context-dependent risks that children may encounter when viewing AIGC videos. To address this gap, we study child-oriented AIGC video reviewing, making three contributions. First, we construct CAVSR, a benchmark of 605 real-world videos collected from multiple platforms, and develop a hierarchical risk taxonomy comprising 6 top-level categories and 26 fine-grained labels to support systematic evaluation of children's viewing risks. Second, we propose QVRS-E, a knowledge- and experience-augmented video reviewing framework that combines multi-agent collaboration with expert and experiential knowledge to support targeted evidence acquisition and fact-grounded reviewing decisions. Third, extensive experiments demonstrate that our method significantly enhances the reviewing of child-related risks integrated with vision-language models, and yields more robust review reports.