🤖 AI Summary
This study investigates the longitudinal educational impact of large language models (LLMs) in software engineering project-based learning (PBL), focusing on their dual role in shaping learning equity and performance differentiation. Employing a two-year longitudinal cohort design, we integrate quantitative performance analysis with qualitative usage observation to compare the pedagogical efficacy of free (earlier-generation) versus paid (state-of-the-art) LLMs in authentic course projects. Our findings provide the first empirical evidence that LLMs function simultaneously as both “equalizers” and “amplifiers”: they significantly enhance overall practical competency—particularly elevating the average performance of students with weak programming foundations—yet concurrently widen the absolute performance gap between high-achieving and average-performing students. These results offer critical empirical grounding and novel theoretical insights for ethically integrating LLMs into SE education, designing adaptive pedagogical strategies, and informing equitable AI-in-education policy.
📝 Abstract
As LLMs reshape software development, integrating LLM-augmented practices into SE education has become imperative. While existing studies explore LLMs' educational use in introductory programming or isolated SE tasks, their impact in more open-ended Project-Based Learning (PBL) remains unexplored. This paper introduces a two-year longitudinal study comparing a 2024 (using early free LLMs, $n$=48) and 2025 (using the latest paid LLMs, $n$=46) cohort. Our findings suggest the latest powerful LLMs' dual role: they act as "equalizers," boosting average performance even for programming-weak students, providing opportunities for more authentic SE practices; yet also as "amplifiers," dramatically widening absolute performance gaps, creating new pedagogical challenges for addressing educational inequities.