Score
Designs and constructs curated datasets of real-world classroom video by collecting and organizing footage, documenting metadata, and ensuring legal and ethical compliance. Develops and applies annotation schemas to label teacher and student behaviors and events, produces training/evaluation splits and quality checks, and prepares the data to capture classroom variability for downstream modeling and analysis.
Educational video design in higher education lacks data-driven optimization tools and open resources, hindering learning effectiveness. To address this, we propose the first open-source, scalable video analytics workflow integrating multimodal feature extraction (frames, audio, transcripts), structured metadata modeling, and supervised machine learning—including regression and feature importance analysis—to enable evidence-informed design iteration. We release the first community-curated, open database of educational video attributes. Empirical validation across two engineering courses identified key pedagogical factors—such as lecture pacing and visual complexity—and informed 12 pedagogical experiments and three international collaborative projects. Our framework establishes a reproducible, generalizable paradigm for optimizing educational video design through empirical, multimodal analysis.
Traditional classroom observation methods suffer from high subjectivity and limited scalability, lacking objective means to assess student attentiveness. To address this gap, this study introduces BAV-Classroom, the first fine-grained video dataset of classroom behaviors collected in Vietnamese higher education institutions, annotated with nine distinct student behavior categories. The work systematically evaluates the performance of YOLO-family models for automated behavior recognition, demonstrating that YOLOv11 achieves superior accuracy and efficiency on this task. Using this model, the study reveals a significant decline in student focus during the latter segments of lectures. This research provides a reliable, data-driven tool for monitoring teaching quality and evaluating student engagement in real-world classroom settings.
To address the critical bottleneck of scarce large-scale, high-quality, real-world classroom behavioral annotation datasets in educational AI research, this paper introduces SCB-dataset—the first fine-grained student behavior detection dataset specifically designed for authentic classroom settings. It comprises 4,003 images with 11,248 precise bounding-box annotations, emphasizing pedagogically significant interactive behaviors such as hand-raising. The dataset is open-source, rigorously annotated, and grounded in real classroom environments, thereby filling a longstanding gap in publicly available benchmarks for classroom behavior analysis. We conduct benchmark experiments using YOLOv7, achieving an mAP of 85.3% on SCB-dataset, empirically validating its quality and utility. SCB-dataset serves as a foundational resource for intelligent classroom behavior understanding and training of education-oriented large language models.
This study addresses the high cost and limited scalability of manual annotation for classroom interactions by proposing a multimodal AI framework for fine-grained, automated recognition of teaching activities (24 classes) and teacher-student discourse (19 classes). Methodologically, it introduces parallel, modality-specific pipelines for video and text processing, integrating contextual window modeling, class-balanced sampling, and multi-label threshold optimization. The framework employs fine-tuned vision-language models, self-supervised video Transformers, and contextualized Transformer classifiers, with ablation against zero-shot large language model prompting. Results show that fine-tuned models significantly outperform prompting: macro-F1 scores reach 0.577 (video) and 0.460 (text), demonstrating the feasibility of scalable, automated classroom feedback systems. The core contribution is the first end-to-end, multimodal, fine-grained recognition framework tailored to instructional settings, coupled with robust training strategies for imbalanced, multimodal educational data.
Fine-grained student behavior analysis in educational settings is hindered by the absence of realistic, multi-label action datasets captured in authentic classroom environments. To address this gap, we introduce SAV—the first large-scale, multi-label student action video dataset curated from real classrooms—comprising 4,324 annotated video clips spanning 15 distinct action classes and explicitly capturing challenging conditions including small objects, high subject density, and severe occlusion. We further propose an education-optimized visual Transformer baseline that integrates fine-grained local attention with spatiotemporal modeling to effectively resolve subtle action discrimination and dense interaction recognition. Evaluated on SAV, our model achieves a mean average precision (mAP) of 67.9%, substantially outperforming existing methods. Both the dataset and source code are publicly released to foster reproducible research in educational behavioral analytics.
This study addresses the lack of structured annotation benchmarks in existing classroom videos for evaluating multimodal models. The authors construct a multimodal teaching observation benchmark comprising 30 international lecture videos segmented into 5,158 fifteen-second clips, annotated with 39 binary-coded visual and non-visual dimensions. For the first time, this benchmark integrates fine-grained scene-level labels with whole-lesson expert ratings and qualitative assessments, establishing a two-tier human reference framework. They propose a Krippendorff’s alpha–based approach to construct reliability- and prevalence-aware labels and evaluate five state-of-the-art vision-language foundation models across tasks involving text-only, text-plus-image-frame, and full-lesson comprehension, using human annotations, expert scoring, and an LLM-as-judge protocol. Results reveal no single model dominates across all tasks; incorporating intermediate frames improves both true and false attribution accuracy, yet models tend to overrate instruction that is procedurally clear but lacks depth—highlighting the irreplaceable role of expert judgment in complex pedagogical assessment.
Existing long-form video generation methods struggle to maintain knowledge consistency and pedagogical narrative coherence in multi-shot STEM instructional videos. This work proposes EduStory, a unified framework that introduces, for the first time, a teaching-state modeling mechanism coupled with script-guided structured narrative control to generate multi-shot videos that are both factually accurate and logically coherent. The core contributions include the construction of EduVideoBench, a diagnostic benchmark featuring multi-granularity annotations; the design of learning-oriented metrics for assessing knowledge fidelity; and a significant reduction in narrative discontinuities, thereby enhancing alignment between generated videos and intended instructional goals. The approach demonstrates substantial progress in both the accuracy of knowledge transmission and the controllability of narrative structure.
This work addresses the lack of systematic evaluation of educational efficacy in existing video generation models, which predominantly focus on perceptual quality or general safety. The authors propose EduVideoBench—the first benchmark for evaluating educational video generation grounded in the Knowledge-Skills-Attitudes (KSA) framework from educational theory. By integrating KSA into generative model assessment, they establish a multidimensional, structured evaluation system for instructional appropriateness and educational safety. Through expert review and qualitative analysis, five state-of-the-art models are systematically assessed, revealing significant deficiencies in knowledge accuracy, skill demonstration, and attitudinal appropriateness. The findings indicate that misalignment in any single KSA dimension can render the generated content educationally ineffective, underscoring a substantial gap between current models and practical classroom deployment.
This work addresses the high barrier, static nature, and limited evolvability of specialized video training datasets by proposing a configurable, self-evolving video data operating system. The system enables users to issue requests via natural language and parameters, automatically optimizing queries and executing parallel retrieval of real videos alongside controllable synthesis to produce domain-specific data packages enriched with full provenance, multidimensional metadata, and reproducible notebooks. Built upon the Model-Context Protocol (MCP), it establishes a dynamic data ecosystem that supports community contributions, governance-driven continuous updates, and flexible “cooking” mechanisms. Experiments demonstrate that this approach substantially reduces dataset construction costs and enhances the training efficiency and iterative capability of multimodal large models in vertical domains.