Is Solving Better Than Evaluating GenAI Solutions?

📅 2026-07-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether evaluating generative artificial intelligence (GenAI)-produced solutions fosters greater learning gains than traditional problem-solving in an advanced algorithms course. Employing a randomized crossover experimental design, the research systematically compares the “solve directly” and “evaluate GenAI solutions” approaches through integrated analyses of student performance, survey responses, and assignment content alignment, focusing on conceptual understanding and transfer ability. As the first study to examine GenAI solution evaluation in a theory-intensive upper-level course, it finds that the evaluation group achieved significantly higher assignment scores, though no significant differences emerged in midterm, final, or overall course grades. Notably, students reported perceiving greater benefit from the evaluation task only when they actively adapted their learning strategies, suggesting a potential mechanism through which such tasks influence strategic effort allocation and metacognitive engagement.
📝 Abstract
As Generative AI (GenAI) tools become increasingly capable of generating solutions to computing assignments, the computing education community is exploring pedagogical approaches that emphasize solution evaluation, verification, and critique alongside traditional solution generation. However, evidence regarding the impact of such evaluation-centered tasks on student learning remains limited, particularly in upper-division, theory-heavy courses. We conducted a randomized A/B crossover study (N=220) in a junior-level algorithms course to compare evaluating GenAI-generated solutions with traditional problem solving. Across six assignments, student working groups either solved challenging algorithmic problems directly or evaluated often-flawed GenAI-generated solutions, with roles reversed midway through the semester. We found no statistically significant differences between groups in midterm scores, final exam scores, overall course grades, or exam problems structurally aligned with the homework interventions. Students received significantly higher homework scores when evaluating GenAI-generated solutions, but this localized advantage did not translate into downstream summative gains. Survey data further indicated that most students reported no change in study habits in response to the intervention; however, those who reported adapting their study strategies rated the GenAI-evaluation assignments as significantly more helpful. These findings suggest that GenAI evaluation redistributes student effort from open-ended solution construction toward verification, diagnosis, and judgment, but does not automatically produce stronger conceptual transfer. We conclude that GenAI-evaluation activities can be incorporated into algorithms coursework without broad performance losses, but meaningful learning gains may require deliberate scaffolding that pushes students beyond simple error diagnosis.
Problem

Research questions and friction points this paper is trying to address.

Generative AI
computing education
solution evaluation
student learning
algorithms course
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generative AI
solution evaluation
computing education
algorithmic problem solving
randomized crossover study