arXiv · 2609.32460
REMEDY: How Far Is Video Generation from Medical Education World Models?
Abstract
Recent video generation models produce realistic videos and show potential as a foundation for world models. These advances create opportunities for generating medical teaching demonstrations, which requires both convincing visual quality and precise procedural actions. However, whether current generators can meet these requirements has not been measured. To address this problem, we introduce Readiness Evaluation of Medical Education Demonstration sYnthesis (REMEDY), to our knowledge, the first benchmark for AI-generated medical teaching demonstrations. REMEDY provides 900 first frames from real demonstration videos, covering 12 tasks across four scenarios: operating room, imaging, clinic and bedside, and resuscitation. Five contemporary open-source video generation models produce 4,500 videos from these frames. We combine task-specific clinical checklists with video and motion quality metrics. Evaluation covers four dimensions: clinical action following, clinical profiles, video quality, and motion quality. Our results show that realistic appearance and temporal consistency do not ensure correct clinical actions. Even the most advanced MiniMax-H3 achieves only 28.25% on strict clinical success rate, and fine-grained clinical actions remain challenging. These findings establish a foundation and roadmap for developing future medical education world models.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Lixing Tan, Yanghao Zhou, Qing Xia, Yuting Guo, Shuai Li, Aimin Hao. 2026-09-26. REMEDY: How Far Is Video Generation from Medical Education World Models?. https://arxiv.org/abs/2609.32460
Cite the original work for its findings. Save a collection to share your selection of sources.