🤖 AI Summary
This work addresses the limitations of existing text-to-image methods in handling long-form script narratives and maintaining visual consistency across multi-page electronic theater programs (ETPs). To overcome these challenges, we propose a multi-agent collaborative framework that directly generates high-quality, stylistically coherent multi-page ETPs from raw scripts while enabling character animation and real-time voice interaction. The framework incorporates a global style anchoring mechanism to ensure cross-page visual consistency and integrates modules for semantic script analysis, hierarchical character composition, portrait animation, customized voice synthesis, and a personalized large language model, thereby transforming static programs into immersive interactive experiences. Extensive experiments on our newly curated, professional-grade ETP-Pro benchmark demonstrate that our approach significantly outperforms current methods in semantic fidelity, aesthetic consistency, and interactivity.
📝 Abstract
Electronic Theater Programs (ETPs) serve as critical promotional media in the performing arts, comprising a multi-page collection of heterogeneous visual assets such as theatrical posters, performance details, and character portraits. However, existing text-to-image paradigms struggle with such complex design tasks due to their inability to comprehend long-context narratives and maintain visual consistency across multiple distinct pages. To address this, we introduce ETPDesigner, a collaborative Multi-Agent framework that directly synthesizes high-quality ETPs from raw dramatic scripts. Emulating a professional design pipeline, our framework orchestrates specialized agents for semantic script analysis, core poster synthesis, functional background generation, and the stratified composition of character assets. Central to ETPDesigner is a global style anchor mechanism that extracts visual priors from the core poster to enforce strict aesthetic uniformity across all generated components. Furthermore, we elevate the ETP from a static publication to an immersive interactive companion. By integrating portrait animation, customized speech synthesis, and persona-grounded Large Language Models (LLMs), our system enables users to engage in real-time, voice-enabled conversations with the generated virtual characters. To rigorously benchmark this task, we construct ETP-Pro, a domain-specific benchmark of professional theater posters and high-quality character portraits. Extensive evaluations demonstrate our method's superiority in producing semantically faithful, aesthetically consistent, and highly interactive program sets.