🤖 AI Summary
This study addresses the discrepancy between visual realism and clinical action accuracy in video generation models for medical education, proposing the first world-model-readiness evaluation framework tailored to this domain. By constructing a benchmark dataset encompassing four scenario categories and twelve tasks, videos were generated using five open-source models and systematically evaluated against clinical checklists and multidimensional quality metrics. The findings reveal that visual fidelity does not guarantee procedural correctness; even the state-of-the-art model achieves a strict clinical success rate of only 28.25%, demonstrating that fine-grained clinical action generation remains a core challenge. This work establishes a rigorous evaluation baseline and provides a roadmap for future research and development in medically grounded video generation.
📝 Abstract
Recent video generation models produce realistic videos and show potential as a foundation for world models. These advances create opportunities for generating medical teaching demonstrations, which requires both convincing visual quality and precise procedural actions. However, whether current generators can meet these requirements has not been measured. To address this problem, we introduce Readiness Evaluation of Medical Education Demonstration sYnthesis (REMEDY), to our knowledge, the first benchmark for AI-generated medical teaching demonstrations. REMEDY provides 900 first frames from real demonstration videos, covering 12 tasks across four scenarios: operating room, imaging, clinic and bedside, and resuscitation. Five contemporary open-source video generation models produce 4,500 videos from these frames. We combine task-specific clinical checklists with video and motion quality metrics. Evaluation covers four dimensions: clinical action following, clinical profiles, video quality, and motion quality. Our results show that realistic appearance and temporal consistency do not ensure correct clinical actions. Even the most advanced MiniMax-H3 achieves only 28.25% on strict clinical success rate, and fine-grained clinical actions remain challenging. These findings establish a foundation and roadmap for developing future medical education world models.