educational assessment

Designing instruments and evaluation protocols to measure learning outcomes and instructional impact (e.g., immediate conceptual gains, next-attempt success) and to compare novel interactive interventions against standard pen-and-paper or lecture-only methods.

educationalassessment

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Traditional lecture-based instruction struggles to help engineering students connect spatial reasoning with abstract mechanical concepts, thereby limiting their conceptual understanding. This study addresses this challenge through a randomized controlled trial in an undergraduate solid mechanics course, directly comparing mixed reality (MR) and physical manipulatives as post-lecture interventions within an authentic STEM instructional setting—the first such comparison to date. A multidimensional evaluation encompassing learning gains, usability, and user experience reveals that both interactive approaches significantly outperform lecture-only instruction. The MR group achieved the highest normalized learning gain (g = 0.57), while the physical manipulatives group reported superior usability and greater learning confidence. Notably, the MR intervention demonstrated good tolerability among students.

abstract representationsconceptual learninginteractive learning environments

The impact of gamification on learning outcomes: experiences from a Biomedical Engineering course

Sep 07, 2025
GR
Gonzalo R. Ríos-Muñoz
🏛️ Universidad Carlos III de Madrid | Instituto de Investigación Sanitaria Gregorio Marañón

Project-based learning (PjBL) in biomedical engineering education suffers from inefficient collaboration, opaque process tracking, and subjective assessment. Method: This study designed and empirically validated a digitally integrated pedagogical framework featuring real-time experimental progress monitoring via a learning platform, embedded structured peer-assessment rubrics, and multi-role collaborative editing with versioned audit trails. A mixed-methods evaluation was conducted using student surveys, project artifact analysis, and instructor feedback. Results: The intervention improved team collaboration quality by 23%, enhanced assessment fairness (inter-rater reliability increased by 0.31 in Krippendorff’s α), boosted instructor process-monitoring efficiency by 40%, and concurrently strengthened student deep engagement and metacognitive capacity. The study’s key contribution lies in pioneering the first systematic integration of a “triple-digital-support” framework—process visualization, bias-mitigated evaluation, and structured collaboration—within advanced biomedical engineering practice instruction, offering a replicable methodology and empirical evidence for digital transformation of PjBL in STEM education.

Enhancing collaboration, transparency, and assessment fairnessIntegrating digital tools in project-based learningStructured learning environment for Biomedical Engineering

This paper addresses the distortion of effect size estimates in educational and psychological intervention research due to differential item functioning (DIF). Moving beyond conventional differential test functioning (DTF) analyses that rely on total-score differences, we propose a novel causal robustness framework grounded in item response theory (IRT). We formally define “impact” as between-group differences in the latent trait distribution and develop a Hausman-type test that integrates DIF modeling directly into causal effect identification—thereby disentangling true construct-level impact from item-specific bias. Methodologically, we introduce a DIF-robust doubly robust estimator and a testable framework for effect generalizability inference. Empirical validation across item-level data from 34 randomized trials shows that DIF correction substantially reduces discrepancies between effect estimates derived from researcher-developed versus independent measures, thereby enhancing construct validity and cross-measure comparability of effect interpretations.

Compares latent trait distribution differences between respondent groupsDevelops robust scaling method for consistent impact estimationProposes effect size for DIF's impact on group comparisons

The Right Kind of Help: Evaluating the Effectiveness of Intervention Methods in Elementary-Level Visual Programming

Dec 12, 2025
AG
Ahana Ghosh
🏛️ Max Planck Institute for Software Systems | Tallinn University

This study investigates the efficacy of three interventions—code-editing recommendations, editing-based quizzes, and metacognitive-strategy quizzes—in elementary visual programming learning, evaluated across both in-learning and post-learning phases. Employing a large-scale empirical design on the Hour of Code: Maze Challenge platform, multimodal data—including behavioral logs, pre/post-tests, and self-report scales—were collected and analyzed. For the first time, this work systematically disentangles and separately assesses intervention effects on immediate learning versus subsequent transfer performance. Results indicate that quiz-based interventions significantly enhance far-transfer outcomes without compromising problem-solving ability; all intervention groups outperformed the control group, with notable improvements in engagement and perceived skill growth. The core contribution lies in establishing the “post-learning phase” as a critical evaluation dimension and empirically validating the specificity of quiz-based interventions in promoting far transfer.

Assesses impact on performance, problem-solving, and student engagementCompares effectiveness during learning and post-learning phasesEvaluates intervention methods in elementary visual programming

This study identifies and empirically validates the “Interaction–Effectiveness Paradox”: although large language models (LLMs) enable richer, metacognitively aware interactions—such as deep knowledge articulation and reflective self-monitoring—compared to search engines, they do not yield statistically significant improvements in learning outcomes. Method: A controlled experiment (N = 20) compared an LLM-based dialogue system with a conventional search interface across authentic learning tasks, integrating qualitative interaction analysis with quantitative learning assessments. Results: While LLMs enhanced interaction quality, they failed to produce a statistically significant gain in overall learning effectiveness. Contribution: This work formally defines the paradox for the first time, revealing that increased interactivity may redistribute—rather than augment—cognitive effort. It advocates for educational AI design that scaffolds, rather than supplants, learners’ active cognitive engagement, offering a novel theoretical framework and practical implications for AI-augmented learning.

Comparing learning effectiveness between LLM dialogue and search interfacesExploring how student cognitive shifts affect educational technology designInvestigating why richer LLM interactions do not improve learning outcomes

Latest Papers

What's happening recently
View more

This study addresses the decline in student engagement and feedback quality often caused by frequent survey requests in traditional course evaluation systems. It proposes a lightweight, randomized weekly feedback approach—termed HRCF—that balances timeliness and depth by administering brief questionnaires to each student during randomly selected weeks. Evaluated over four years across 103 courses and 24,216 students in authentic large-scale teaching settings, regression analyses reveal that sustained implementation of HRCF for one semester yields an average increase of 0.045–0.048 points in learning-related end-of-term evaluation scores for small- to medium-sized courses. However, no statistically significant effects were observed for large-enrollment courses or metrics related to instructional organization.

course experienceeducational evaluationfeedback quality

This study investigates the generalizability of data-driven approaches to reconstructing intelligent tutoring systems in non-preselected instructional units. Focusing on four middle school mathematics topics that had not undergone prior efficacy screening, the authors implemented a system redesign grounded in data-driven instructional design and learning behavior analysis, followed by a classroom-based randomized controlled trial to evaluate its impact. Although the intervention did not yield statistically significant gains in overall learning outcomes, it led to significant improvements in students’ engaged learning time, volume of skill practice, and breadth of knowledge acquisition. This work represents the first validation of data-driven reconstruction methods in uncurated instructional contexts, thereby extending their applicability to more authentic and diverse educational settings.

classroom studydata-driven redesigneffectiveness

This study addresses the challenge of developing fine motor skills in K–2 remote art education, where traditional activities prove ineffective and teachers lack real-time feedback mechanisms. To bridge this gap, the authors propose Chameleon Clippers—an interactive pair of scissors embedded with sensors that leverages a tangible user interface and digital feedback to provide immediate guidance as students cut along lines. Designed as a low-cost augmentation of conventional classroom tools, this system represents the first exploration of tangible interaction in remote art instruction for young learners. Preliminary findings indicate high student engagement, positive responsiveness to real-time feedback, and an overall sense of enjoyment, demonstrating the tool’s effectiveness in enhancing interactivity and support in remote teaching contexts.

art educationfine motor skillsK-2 learners

This study addresses the limitations of traditional manual grading in high school STEM education, which restricts the frequency of formative assessments and delays feedback. The authors deployed the first fully automated handwriting assessment platform supporting Error Carried Forward (ECF) logic in A-Level science courses, enabling fine-grained, structured feedback on mathematics and science problems. Leveraging handwriting recognition and intelligent scoring algorithms, the system reduced grading turnaround from 11.2 days to under 0.1 days and quadrupled student practice volume. Students in the intervention group demonstrated a statistically significant 16.8% average improvement (p<0.01) in mock exam scores, alongside enhanced self-efficacy and error-correction efficiency, thereby validating the positive impact of high-frequency automated assessment on learning outcomes.

A-levelautomated markingformative assessment

Hot Scholars

XZ

Xiaoming Zhai

Associate Professor, University of Georgia
Science EducationAIAssessment
CB

Conrad Borchers

Carnegie Mellon University
Educational Data MiningLearning AnalyticsIntelligent Tutoring SystemsSelf-Regulated Learning
RC

Ruth Cobos

Universidad Autonoma de Madrid
Learning AnalyticsMachine LearningSentiment AnalysisCSCW
JS

John Stamper

Human-Computer Interaction Institute, Carnegie Mellon University
Artificial IntelligenceEducational Data MiningIntelligent Tutoring Systems
AS

Atsushi Shimada

Professor of Kyushu University
Pattern RecognitionLearning AnalyticsEducational Data MiningEdTech