🤖 AI Summary
This study addresses the high barrier to data video creation and the difficulty of balancing accuracy with narrative coherence in end-to-end generation by proposing an automated framework based on declarative multi-agent orchestration. The method automatically synthesizes charts, narration, and animations from raw tabular data. Its core innovations include DVSpec, a unified specification ensuring precise audio-visual alignment; a "generate-then-orchestrate" strategy that preserves global narrative consistency; and the integration of large language models to support human-in-the-loop interaction. Experimental results demonstrate that the system achieves an execution success rate exceeding 95%, improves generation quality by 83%, and enhances creative efficiency by nearly 80%, significantly reducing users' cognitive load.
📝 Abstract
Data videos communicate data insights through dynamic charts, voice narration, and synchronized animations, and have become a widely adopted form of data storytelling. However, producing them requires expertise in data analysis, narrative design, and video editing. Static visualization tools lack narrative and animation capabilities; authoring tools rely on pre-prepared charts rather than raw data; and pixel-level models generate videos end-to-end but cannot guarantee data accuracy or provenance. End-to-end automatic generation faces two core challenges: how to uniformly represent charts, narration, and animations together with their temporal relationships, and how to efficiently search a vast design space for narrative-coherent compositions. We present DataMagic, which authors data videos from raw tabular data through declarative multi-agent orchestration. First, the declarative specification DVSpec unifies charts, narration, and animations with data-bound references and declarative synchronization, ensuring data provenance and automatic audio-visual alignment. Second, a"Generate-then-Orchestrate"multi-agent strategy generates candidate scenes in parallel and then optimizes narrative coherence through global orchestration. DVSpec provides a shared state for three complementary interaction modes, bridging full automation with fine-grained human control. Evaluations on 109 real-world samples show that even the most advanced LLM (e.g., GPT-5) achieves only 2.13/5 with execution success rates between 48.62% and 86.24%; DataMagic improves quality to 3.89 (+83%) with success rates above 95%, with the most significant gains in animation and narrative dimensions. A user study shows that, compared to a conversational LLM workflow, DataMagic improves creation efficiency (79.7% reduction in task time) and reduces perceived cognitive load. Project page: https://github.com/HKUSTDial/DataMagic.