classroom video dataset curation

Designs and constructs curated datasets of real-world classroom video by collecting and organizing footage, documenting metadata, and ensuring legal and ethical compliance. Develops and applies annotation schemas to label teacher and student behaviors and events, produces training/evaluation splits and quality checks, and prepares the data to capture classroom variability for downstream modeling and analysis.

classroomvideodatasetcuration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.03
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Educational video design in higher education lacks data-driven optimization tools and open resources, hindering learning effectiveness. To address this, we propose the first open-source, scalable video analytics workflow integrating multimodal feature extraction (frames, audio, transcripts), structured metadata modeling, and supervised machine learning—including regression and feature importance analysis—to enable evidence-informed design iteration. We release the first community-curated, open database of educational video attributes. Empirical validation across two engineering courses identified key pedagogical factors—such as lecture pacing and visual complexity—and informed 12 pedagogical experiments and three international collaborative projects. Our framework establishes a reproducible, generalizable paradigm for optimizing educational video design through empirical, multimodal analysis.

Absence of open datasets to analyze video design impact on engagementLack of scalable tools for data-driven educational video design improvementsNeed for evidence-based workflow to enhance educational video effectiveness

Traditional classroom observation methods suffer from high subjectivity and limited scalability, lacking objective means to assess student attentiveness. To address this gap, this study introduces BAV-Classroom, the first fine-grained video dataset of classroom behaviors collected in Vietnamese higher education institutions, annotated with nine distinct student behavior categories. The work systematically evaluates the performance of YOLO-family models for automated behavior recognition, demonstrating that YOLOv11 achieves superior accuracy and efficiency on this task. Using this model, the study reveals a significant decline in student focus during the latter segments of lectures. This research provides a reliable, data-driven tool for monitoring teaching quality and evaluating student engagement in real-world classroom settings.

automated monitoringclassroom behavior monitoringcomputer vision

SCB-dataset: A Dataset for Detecting Student Classroom Behavior

Apr 05, 2023
YF
Yang Fan
🏛️ Jinan University

To address the critical bottleneck of scarce large-scale, high-quality, real-world classroom behavioral annotation datasets in educational AI research, this paper introduces SCB-dataset—the first fine-grained student behavior detection dataset specifically designed for authentic classroom settings. It comprises 4,003 images with 11,248 precise bounding-box annotations, emphasizing pedagogically significant interactive behaviors such as hand-raising. The dataset is open-source, rigorously annotated, and grounded in real classroom environments, thereby filling a longstanding gap in publicly available benchmarks for classroom behavior analysis. We conduct benchmark experiments using YOLOv7, achieving an mAP of 85.3% on SCB-dataset, empirically validating its quality and utility. SCB-dataset serves as a foundational resource for intelligent classroom behavior understanding and training of education-oriented large language models.

Classroom Behavior AnalysisDeep Learning in EducationLarge-scale Datasets

Exploring Automated Recognition of Instructional Activity and Discourse from Multimodal Classroom Data

Nov 26, 2025
IB
Ivo Bueno
🏛️ Technical University of Munich | University of Tübingen | Gordon College | State University of New York at Albany | University of Virginia

This study addresses the high cost and limited scalability of manual annotation for classroom interactions by proposing a multimodal AI framework for fine-grained, automated recognition of teaching activities (24 classes) and teacher-student discourse (19 classes). Methodologically, it introduces parallel, modality-specific pipelines for video and text processing, integrating contextual window modeling, class-balanced sampling, and multi-label threshold optimization. The framework employs fine-tuned vision-language models, self-supervised video Transformers, and contextualized Transformer classifiers, with ablation against zero-shot large language model prompting. Results show that fine-tuned models significantly outperform prompting: macro-F1 scores reach 0.577 (video) and 0.460 (text), demonstrating the feasibility of scalable, automated classroom feedback systems. The core contribution is the first end-to-end, multimodal, fine-grained recognition framework tailored to instructional settings, coupled with robust training strategies for imbalanced, multimodal educational data.

Automates recognition of discourse patterns from classroom transcriptsAutomates recognition of instructional activities from classroom videoEstablishes scalable AI foundation for teacher feedback systems

Towards Student Actions in Classroom Scenes: New Dataset and Baseline

Sep 02, 2024
ZT
Zhuolin Tan
🏛️ Chongqing University of Posts and Telecommunications

Fine-grained student behavior analysis in educational settings is hindered by the absence of realistic, multi-label action datasets captured in authentic classroom environments. To address this gap, we introduce SAV—the first large-scale, multi-label student action video dataset curated from real classrooms—comprising 4,324 annotated video clips spanning 15 distinct action classes and explicitly capturing challenging conditions including small objects, high subject density, and severe occlusion. We further propose an education-optimized visual Transformer baseline that integrates fine-grained local attention with spatiotemporal modeling to effectively resolve subtle action discrimination and dense interaction recognition. Evaluated on SAV, our model achieves a mean average precision (mAP) of 67.9%, substantially outperforming existing methods. Both the dataset and source code are publicly released to foster reproducible research in educational behavioral analytics.

Challenges in detecting nuanced actions in classroom settings.Lack of accessible datasets for student action analysis.Need for advanced methods to improve action detection accuracy.

Latest Papers

What's happening recently
View more

This study addresses the lack of structured annotation benchmarks in existing classroom videos for evaluating multimodal models. The authors construct a multimodal teaching observation benchmark comprising 30 international lecture videos segmented into 5,158 fifteen-second clips, annotated with 39 binary-coded visual and non-visual dimensions. For the first time, this benchmark integrates fine-grained scene-level labels with whole-lesson expert ratings and qualitative assessments, establishing a two-tier human reference framework. They propose a Krippendorff’s alpha–based approach to construct reliability- and prevalence-aware labels and evaluate five state-of-the-art vision-language foundation models across tasks involving text-only, text-plus-image-frame, and full-lesson comprehension, using human annotations, expert scoring, and an LLM-as-judge protocol. Results reveal no single model dominates across all tasks; incorporating intermediate frames improves both true and false attribution accuracy, yet models tend to overrate instruction that is procedurally clear but lacks depth—highlighting the irreplaceable role of expert judgment in complex pedagogical assessment.

benchmarkclassroom videosmodel evaluation

Existing long-form video generation methods struggle to maintain knowledge consistency and pedagogical narrative coherence in multi-shot STEM instructional videos. This work proposes EduStory, a unified framework that introduces, for the first time, a teaching-state modeling mechanism coupled with script-guided structured narrative control to generate multi-shot videos that are both factually accurate and logically coherent. The core contributions include the construction of EduVideoBench, a diagnostic benchmark featuring multi-granularity annotations; the design of learning-oriented metrics for assessing knowledge fidelity; and a significant reduction in narrative discontinuities, thereby enhancing alignment between generated videos and intended instructional goals. The approach demonstrates substantial progress in both the accuracy of knowledge transmission and the controllability of narrative structure.

instructional video generationknowledge consistencymulti-shot video

This work addresses the lack of systematic evaluation of educational efficacy in existing video generation models, which predominantly focus on perceptual quality or general safety. The authors propose EduVideoBench—the first benchmark for evaluating educational video generation grounded in the Knowledge-Skills-Attitudes (KSA) framework from educational theory. By integrating KSA into generative model assessment, they establish a multidimensional, structured evaluation system for instructional appropriateness and educational safety. Through expert review and qualitative analysis, five state-of-the-art models are systematically assessed, revealing significant deficiencies in knowledge accuracy, skill demonstration, and attitudinal appropriateness. The findings indicate that misalignment in any single KSA dimension can render the generated content educationally ineffective, underscoring a substantial gap between current models and practical classroom deployment.

benchmarkeducational safetyeducational validity

This work addresses the high barrier, static nature, and limited evolvability of specialized video training datasets by proposing a configurable, self-evolving video data operating system. The system enables users to issue requests via natural language and parameters, automatically optimizing queries and executing parallel retrieval of real videos alongside controllable synthesis to produce domain-specific data packages enriched with full provenance, multidimensional metadata, and reproducible notebooks. Built upon the Model-Context Protocol (MCP), it establishes a dynamic data ecosystem that supports community contributions, governance-driven continuous updates, and flexible “cooking” mechanisms. Experiments demonstrate that this approach substantially reduces dataset construction costs and enhances the training efficiency and iterative capability of multimodal large models in vertical domains.

data infrastructuredataset evolutiondomain-specific data

Hot Scholars

KC

Kehai Chen

Harbin Institute of Technolgy (Shenzhen)
LLMNatural Language ProcessingAgentMulti-model Generation
WW

Wenxuan Wang

Renmin University of China
AI SafetyTrustworthy AILarge Language ModelsEvaluation
RJ

Rahul Jain

PhD Student, Elmore Family School of Electrical and Computer Engineering, Purdue University
Deep LearningComputer VisionCausality Graph modelsHuman-Computer Interaction