🤖 AI Summary
This study addresses the narrative fragmentation and inefficient communication of core messages in scientific presentations by proposing a training-free multi-agent framework. Through iterative collaboration encompassing content planning, visual generation, layout optimization, and fact-checking, the framework constructs audience-centric scientific slides. Furthermore, it introduces ConfArena, a novel evaluation framework that simulates conference scenarios to perform page-level detection of defects such as data fabrication and image degradation, achieving automated assessment highly aligned with human preferences. Blind evaluations demonstrate that the proposed method outperforms mainstream open-source and commercial systems with a 77% preference rate, reduces inference token consumption by approximately fourfold, and accurately identifies diverse presentation defects.
📝 Abstract
Scientific presentations are more than summaries of research papers. They need to present the work in a coherent sequence, explain the main ideas clearly, and help the audience follow the presentation. We present SlideLab, a training-free multi-agent framework for generating scientific presentations from research papers. SlideLab first plans the presentation narrative, then builds and iteratively refines a shared slide deck using agents for content planning, visual generation, layout refinement, and grounding verification. In a blind human preference study, SlideLab was preferred over both open-source and commercial systems on 77% of papers while using roughly 4 times fewer inference tokens than the strongest open-source baseline. We also introduce ConfArena, an audience-oriented evaluation framework that simulates a conference room and assesses presentations slide by slide. ConfArena matches human system rankings and detects injected presentation problems, including falsified numbers, degraded figures, dropped slides, and shuffled slide order.