Reimagining Assessment in the Age of Generative AI: Lessons from Open-Book Exams with ChatGPT

๐Ÿ“… 2026-05-12
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
The widespread adoption of generative AI poses significant challenges to traditional academic assessment, as it often fails to authentically reflect studentsโ€™ understanding and competencies. This study addresses this issue by permitting engineering students to use ChatGPT during open-book examinations while requiring submission of their interaction logs. The assessment focus thereby shifts from independent problem-solving to evaluating studentsโ€™ abilities to judge, verify, and engineer prompts for AI-generated solutions. Through qualitative analysis of authentic interaction data, three distinct usage patterns emerged: answer retrieval, collaborative prompting, and critical validation. Findings indicate that students demonstrated the strongest reasoning capabilities when identifying and correcting AI errors. Moreover, transparent integration of AI not only reduced evasive behaviors but also fostered self-regulated, practice-oriented learning, offering a novel paradigm for competency-based assessment in the age of artificial intelligence.
๐Ÿ“ Abstract
Generative AI systems such as ChatGPT challenge traditional assumptions about academic assessment by enabling students to generate explanations, code, and solutions in real time. Rather than attempting to restrict AI use, this study investigates how students actually interact with such systems during formal evaluation. Engineering students were permitted to use ChatGPT during take-home open-book exams and were required to submit interaction transcripts alongside exam solutions. This provided direct observational evidence of reasoning processes rather than relying on self-reported behavior. Qualitative analysis revealed three progressive patterns of use: answer retrieval, guided collaboration, and critical verification. While some students initially copied questions verbatim and received generic responses, many refined prompts iteratively and tested outputs. Some of the strongest evidence of reasoning appeared when students evaluated incorrect or incomplete AI responses, revealing evaluative reasoning through debugging, comparison, and justification. The presence of generative AI shifted the cognitive task of assessment from producing solutions to assessing solution validity. The findings suggest that, in AI-mediated assessment environments, correctness of final answers alone may no longer provide sufficient evidence of comprehension. Instead, competencies such as prompt formulation, verification, and judgment become visible indicators of learning. Transparent integration of AI appeared to reduce focus on rule avoidance and promote self-regulation. Assessments should evolve to evaluate reasoning about solutions rather than independent solution production. Generative AI therefore does not invalidate assessment but has the potential to expose deeper forms of understanding aligned with professional practice.
Problem

Research questions and friction points this paper is trying to address.

generative AI
academic assessment
reasoning
ChatGPT
open-book exams
Innovation

Methods, ideas, or system contributions that make the work stand out.

generative AI
assessment redesign
evaluative reasoning
prompt formulation
AI-mediated learning
๐Ÿ”Ž Similar Papers
No similar papers found.
Q
Qusay H. Mahmoud
Faculty of Engineering and Applied Science, Ontario Tech University, Oshawa, Ontario L1G 0C5 Canada