design educational assessments

Designs, builds, and evaluates assessment instruments and measurement systems used to quantify learner knowledge, skills, or cognitive states in educational settings. This work includes creating rubrics and item banks, specifying scoring and validation procedures, and implementing adaptive-assessment algorithms and real-time inference methods that detect learner state and adjust item difficulty or content sequencing.

designeducationalassessments

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.06
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$196K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the challenges teachers face in assessment design, particularly difficulties in authoring and a lack of tools supporting iterative development. Through a seven-month co-design process with 13 educators, the authors introduce a novel conceptual model that characterizes the dual processes of assessment creation and requirement iteration. Building on this model, they developed Ripplet, a web-based tool leveraging large language models (LLMs) to support assessment design. Ripplet incorporates multi-level reusable interaction mechanisms that facilitate a shift from generative to curatorial practices, thereby encouraging teachers’ reflective engagement with assessment quality. A user study with 15 teachers demonstrated that using Ripplet led to higher-quality formative assessments, more meaningful investment of effort, and the successful completion of assessment tasks previously deemed infeasible.

assessment authoringeducational designformative assessment

Implementation Considerations for Automated AI Grading of Student Work

Jun 09, 2025
ZT
Zewei Tian
🏛️ University of Washington | Hensun Innovation

This study addresses the trustworthy deployment of AI-based automated grading systems in K–12 education—specifically, how to enable efficient and interpretable AI-assisted assessment while preserving teacher pedagogical autonomy and student trust. Method: Drawing on a co-design pilot involving 19 teachers, we integrated educational data mining, behavioral log analysis, structured surveys, and in-depth interviews. Contribution/Results: Findings reveal broad teacher acceptance of AI-generated narrative formative feedback, yet systematic rejection of fully automated scoring. We propose a “teacher-centered, human-AI collaborative, feedback-enhancing (not replacing)” paradigm for trustworthy AI assessment, with clearly delineated human–AI responsibility boundaries. Empirical results demonstrate that AI-augmented feedback significantly increases students’ revision willingness and response speed, whereas fully automated scoring undermines teacher authority and erodes student trust. The study provides empirically grounded design and deployment guidelines for educational AI systems.

Assessing teacher trust in AI-generated rubrics and feedbackBalancing AI efficiency with human oversight in gradingExploring AI grading platform implementation in K-12 classrooms

This study addresses the challenge of accurately tracking students’ dynamic mastery of specific skills when the Q-matrix is unknown. Building upon dynamic cognitive diagnosis models, it compares a joint estimation approach—simultaneously inferring the Q-matrix and learning trajectories—with a two-step strategy that first estimates the Q-matrix and then analyzes skill development. Leveraging reading game data and item text embeddings, the research investigates vocabulary and comprehension growth among second- to third-grade students. The authors propose a bias-corrected two-step method and use simulation studies to delineate the conditions under which each approach performs best: joint modeling proves more reliable when the Q-matrix is uncertain and items vary across grade levels. Empirical results indicate that both methods identify a general trend toward mastering both skills, yet they diverge in estimating the proportion of partial mastery in third grade, underscoring the substantive impact of modeling choices on diagnostic conclusions.

cognitive diagnosisdynamic learningmodel comparison

Assessing the Quality of AI-Generated Exams: A Large-Scale Field Study

Aug 08, 2025
CI
Calvin Isley
🏛️ Harvard University | Stanford University | Microsoft Research | Microsoft

This study addresses the psychometric validity of generative AI for item authoring in educational assessment. We propose the first LLM-based framework for self-critical, iterative item generation: large language models automatically generate test items, which are then refined through multiple rounds of AI-driven evaluation and revision. A large-scale empirical validation was conducted across 91 university classrooms in the U.S. (N ≈ 1700), embedded within authentic instructional settings. Items were rigorously evaluated using Item Response Theory (IRT) to estimate key psychometric properties—including difficulty and discrimination. Results demonstrate no statistically significant differences between AI-generated and expert-authored items on these core metrics. This constitutes the first real-world evidence of psychometric equivalence and practical utility of LLM-authored assessments. The study provides a reproducible methodology and empirical foundation for AI-augmented educational measurement.

Assessing AI-generated questions in real-world educational settingsComparing AI-generated and expert-created exam questionsEvaluating psychometric quality of AI-generated exam questions

STEM reading materials lack lightweight, ready-to-deploy tools for assessing students’ domain-specific background knowledge. Method: This study proposes K-tool, the first system enabling single-text-driven, corpus-free automatic generation of domain vocabulary tests. It identifies the core domain via topic detection, models lexical semantic relationships using word embeddings and co-occurrence features, and automatically selects highly relevant target words and generates semantically plausible distractors. The system supports real-time diagnostic assessment and knowledge-state prediction for middle and high school students. Contribution/Results: The architecture is empirically validated; preliminary experiments demonstrate that generated tests significantly differentiate students across knowledge levels (p < 0.01), indicating potential for just-in-time instructional intervention. Its core innovation lies in decoupling test generation from pre-built corpora, thereby enabling “one-text-one-test” lightweight, deployable automated assessment.

Assessing student readiness to understand STEM domain textsAutomated measurement of student background knowledge for text comprehensionGenerating topical vocabulary tests for specific reading passages

Latest Papers

What's happening recently
View more

In the context of widespread large language model (LLM) deployment, traditional assessments may become invalid due to systematic differences between human and AI response behaviors. This study pioneers the extension of differential item functioning (DIF) analysis—a psychometric technique—into the domain of human–AI ability comparison. By integrating negative contrast analysis with item-total correlation-based discrimination methods, the work establishes a novel assessment design paradigm tailored for the AI era. Leveraging data from high school chemistry diagnostic tests and college entrance examinations, the proposed approach is empirically validated. Expert analyses further identify key task dimensions influencing AI performance, offering both theoretical grounding and practical pathways for developing next-generation educational assessments that are robust against AI misuse while ensuring fairness and validity.

AI misuseassessment designdifferential item functioning

This study addresses the challenge of effectively evaluating the adaptive personalization of educational reading materials in the absence of large-scale real learner data. The authors propose a theory-driven simulated learner framework that integrates, for the first time, the Construction-Integration memory model, DIME reader characteristics, the KREC misconception correction mechanism, and the New Dale-Chall readability metric, further incorporating knowledge ontologies and Bayesian Knowledge Tracing (BKT) to enable dynamic content adaptation. This approach allows for the evaluation of adaptive reading systems without requiring human participants. Experimental results demonstrate significantly improved learning outcomes in computer science, a small positive (though statistically non-significant) trend in inorganic chemistry, and neutral to slightly negative effects in general biology.

adaptive personalizationeducational readingslearning outcomes

The widespread adoption of generative AI poses significant challenges to traditional academic assessment, as it often fails to authentically reflect students’ understanding and competencies. This study addresses this issue by permitting engineering students to use ChatGPT during open-book examinations while requiring submission of their interaction logs. The assessment focus thereby shifts from independent problem-solving to evaluating students’ abilities to judge, verify, and engineer prompts for AI-generated solutions. Through qualitative analysis of authentic interaction data, three distinct usage patterns emerged: answer retrieval, collaborative prompting, and critical validation. Findings indicate that students demonstrated the strongest reasoning capabilities when identifying and correcting AI errors. Moreover, transparent integration of AI not only reduced evasive behaviors but also fostered self-regulated, practice-oriented learning, offering a novel paradigm for competency-based assessment in the age of artificial intelligence.

academic assessmentChatGPTgenerative AI

This study addresses the limitations of traditional static assessments, which overemphasize penalty for errors and offer weak diagnostic insight, as well as oral examinations, whose validity is compromised by performance anxiety and power imbalances. To overcome these issues, the paper proposes an automated, human–AI dialogic “Socratic test” that integrates dynamic assessment, a multimodal interaction space, real-time scaffolding informed by Bloom’s taxonomy, and structured scoring based on the SOLO taxonomy within a non-compensatory cumulative architecture to support mastery-oriented evaluation. The approach formally operationalizes the Zone of Proximal Development (ZPD) through progressive scaffolding and ensures measurement reliability via human–AI alignment. This novel assessment framework significantly enhances the precision of dynamically probing students’ cognitive boundaries, effectively balancing diagnostic utility with fairness.

academic hierarchiesconstruct-irrelevant varianceoral examination

Hot Scholars

PD

Paul Denny

Professor, University of Auckland
Educational technologyComputer Science Education
TB

Tiffany Barnes

Distinguished Professor of Computer Science, North Carolina State University
Educational data miningSerious GamesArtificial IntelligenceBroadening Participation
XZ

Xiaoming Zhai

Associate Professor, University of Georgia
Science EducationAIAssessment
HQ

Huamin Qu

Chair Professor, Hong Kong University of Science and Technology
Data visualizationHuman-Computer InteractionExplainable AIE-Learning
CC

Clayton Cohn

PhD Student, Vanderbilt University
NLPLLMsAIEDAgents