Score
Designs and produces informational and expressive artifacts—such as text, images, audio, video, and multimedia—and the accompanying metadata, formats, and templates required for publication and distribution. Builds and refines editorial workflows, content structures, localization and formatting processes, and optimization strategies to meet audience needs, platform constraints, and measurable communication goals.
研究提出了一种自动视频编辑方法,用于生成场景预览、视频摘要和电影预告片,探讨了自动化与创意性之间的关系。
To address challenges in multimodal narrative storytelling—including poor structural control, weak cross-modal consistency, and coarse-grained editing—this paper proposes a graph-node-based multimodal content generation system. Methodologically, it introduces a node-centric editing framework that maps text, images, audio, and video to editable graph nodes; employs a task-selection agent for dynamic generation orchestration; and integrates multimodal large language models, context-aware generation, and natural language understanding to enable node-level precise editing, automatic parallel narrative branching, and cross-modal co-evolution. Contributions include: (1) the first fine-grained, interpretable, human-in-the-loop multimodal narrative iterative generation system; (2) significant performance gains on story outline generation; and (3) user studies demonstrating a 42% improvement in editing efficiency, enhanced creative flexibility, and validated alignment between controllability and effectiveness.
Existing element-attribute grid representations for graphic design completion tasks struggle to model variable-length, type-heterogeneous, and multimodal (text-image) structures. Method: We propose a unified interleaved multimodal tokenized document model that jointly encodes syntactic and semantic structures of markup languages (e.g., SVG/HTML) alongside variable-size, alpha-channel-aware local image generation. We introduce a specialized image quantizer for efficient transparent-image tokenization and integrate an enhanced code-large language model with an interleaved multimodal sequence architecture. Contribution/Results: Our model achieves significant improvements over baselines on three design completion tasks—missing template attributes, image synthesis, and text generation—demonstrating its effectiveness in jointly modeling structural logic and visual semantics in design documents.
This study investigates the irreplaceable role of human creativity in AI-augmented environments, specifically examining human–AI collaboration mechanisms when AI tools are integrated into news short-video production. Method: A 14-week field experiment was conducted in a student newsroom, employing qualitative observation, multimodal content analysis, and reflective practice to evaluate an AI-driven web-to-video workflow. Contribution/Results: The study introduces the novel theoretical framing of AI as a “creative springboard”—not a substitute—for human editors. It empirically demonstrates that editorial critical thinking is indispensable for error correction, aesthetic judgment, and creative refinement. Leveraging multimodal AI tools, the team efficiently produced high-quality news videos achieving over 500,000 total views. Findings confirm that human editors uniquely compensate for AI’s inherent limitations—such as contextual insensitivity and normative bias—and catalyze innovative, contextually grounded solutions, thereby affirming their essential agency in AI-mediated creative workflows.
This work proposes an end-to-end automated visual asset generation method to address inefficiencies in digital collage creation, including cumbersome retrieval, manual image matting, and disorganized asset management. By integrating image annotation, object detection, and segmentation techniques, the approach uniquely leverages a large language model (LLM) to interpret user-provided narrative descriptions and automatically generate semantically coherent primary and auxiliary tags. These tags guide the multi-scale cropping and semantic-level clustering of visual assets, ensuring alignment with the intended narrative. The proposed framework significantly enhances both the efficiency of asset preparation and the semantic consistency between source materials and the creative narrative, thereby enabling users to focus more intently on artistic composition and expressive intent.
This study addresses the lack of systematic understanding regarding the creation, application, and organizational impact of data visualization style guides. Through interviews with nine authors from journalism, government, and industry, complemented by a cross-case analysis of 26 published guides, the paper proposes the PRISM socio-technical framework to elucidate their operational logic across four dimensions: Purpose, Rules and mechanisms, Institutional enforcers, and Situational flexibility. The findings reveal an inherent tension between standardization and adaptability, demonstrating that publicly available guides represent only partial manifestations of more comprehensive internal systems. By unpacking how these guides function in practice, the research offers both theoretical grounding and novel practical insights for the future development of visualization design standards.
该研究提出Live Artifacts,通过封装生成逻辑来创建动态媒体,使用LiveCanvas系统让创作者在视觉画布中管理动态行为和生成持久性。
This study addresses the lack of systematic understanding regarding creators’ usage patterns of base and fine-tuned models—such as LoRA adapters—in the current open-source image generation ecosystem. The authors construct a large-scale dataset comprising six million generated images along with their associated metadata, enabling the first empirical analysis of how 22.4K base models and 154K LoRA models are combined and utilized in real-world creative workflows. Through data mining, metadata analysis, and log correlation, the research uncovers distinctive strengths and inherent challenges within this ecosystem. These findings provide empirical grounding for enhancing its sustainability and innovation potential, while the publicly released high-quality dataset supports further community engagement and academic inquiry.
This work addresses the challenges posed by the heterogeneous multimodal nature of enterprise policy documents, which often cause large language models to hallucinate, disrupt table structures, and lack end-to-end controllability—resulting in labor-intensive manual processing requiring 2–3 days per document. To overcome these limitations, the authors propose a governed multi-agent collaboration framework grounded in a shared, versioned rule repository. The framework integrates large language models (LLMs), vision-language models (VLMs), schema validation, and human-in-the-loop mechanisms through six specialized agents that collaboratively perform parsing, multimodal extraction, consistency verification, evaluation, iterative refinement, and personalized artifact generation, while ensuring full traceability across the pipeline. Evaluated on 120 real-world documents, the approach achieves a 96% success rate, automatically extracts 3,896 rules (71.4% auto-approved), produces 812 deployable artifacts, and reduces per-document processing time to 40–125 minutes.
Existing procedural material generation methods merely replicate node graph structures without capturing the underlying design logic employed by experts, often yielding suboptimal results. This work proposes a process-driven generation paradigm that, for the first time, treats expert creation processes as first-class representations. By automatically analyzing tutorial videos, the approach extracts textualized process trajectories that encode design steps, parameter settings, and intent. Leveraging pretrained large language models, it constructs a ProcessSynthesizer and a Compiler to generate user-aligned trajectories and compile them into executable Blender material graphs. Expert evaluations demonstrate that the generated materials better reflect professional design strategies and require fewer edits, while a user study with 150 participants confirms significant improvements over existing systems in both output quality and editing efficiency.