๐ค AI Summary
Current nursing education relies on subjective, low-efficiency instructor feedback, limiting scalability and standardization of training. To address this, we propose the first video-language model (VLM) framework tailored for clinical skills assessment. Our method innovatively integrates curriculum learning, fine-grained action decomposition, and temporal localization to enable progressive analysisโfrom action recognition to procedural reasoning. Specifically, it unifies high-level action recognition, step-order modeling, and natural language inference to support error diagnosis, temporal deviation localization, and generation of interpretable textual feedback. Evaluated on synthetic surgical video data, our framework achieves 92.3% accuracy in error detection and sub-1.8-second temporal localization error. This work significantly enhances assessment consistency, scalability, and pedagogical utility, establishing a novel paradigm for AI-driven formative evaluation in clinical education.
๐ Abstract
Consistent high-quality nursing care is essential for patient safety, yet current nursing education depends on subjective, time-intensive instructor feedback in training future nurses, which limits scalability and efficiency in their training, and thus hampers nursing competency when they enter the workforce. In this paper, we introduce a video-language model (VLM) based framework to develop the AI capability of automated procedural assessment and feedback for nursing skills training, with the potential of being integrated into existing training programs. Mimicking human skill acquisition, the framework follows a curriculum-inspired progression, advancing from high-level action recognition, fine-grained subaction decomposition, and ultimately to procedural reasoning. This design supports scalable evaluation by reducing instructor workload while preserving assessment quality. The system provides three core capabilities: 1) diagnosing errors by identifying missing or incorrect subactions in nursing skill instruction videos, 2) generating explainable feedback by clarifying why a step is out of order or omitted, and 3) enabling objective, consistent formative evaluation of procedures. Validation on synthesized videos demonstrates reliable error detection and temporal localization, confirming its potential to handle real-world training variability. By addressing workflow bottlenecks and supporting large-scale, standardized evaluation, this work advances AI applications in nursing education, contributing to stronger workforce development and ultimately safer patient care.