Comprehension-Performance Gap in GenAI-Assisted Brownfield Programming: A Replication and Extension

📅 2025-11-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates how generative AI programming assistants (e.g., GitHub Copilot) affect developers’ code comprehension during legacy code maintenance tasks—a critical yet underexplored aspect of AI-augmented software engineering. Method: Using a within-subject experimental design, 18 graduate students completed functional implementation tasks on legacy code; performance was quantified via task completion time, test pass rate, and multidimensional code comprehension metrics. Contribution/Results: Copilot significantly improved productivity—reducing task time by 37% on average and increasing test pass rate by 22%—yet yielded no measurable improvement in code comprehension. Crucially, no significant correlation emerged between comprehension scores and task performance. This study provides the first empirical evidence of a “Comprehension–Performance Gap” in generative AI-assisted programming: AI tools accelerate development without fostering deeper cognitive understanding of code. These findings challenge the implicit assumption that AI-driven efficiency gains entail corresponding improvements in expertise, offering foundational theoretical insights and practical implications for the design, evaluation, and pedagogical integration of AI programming tools.

Technology Category

Natural Language Processing: Code Generation / Program Synthesis from Natural LanguageHumans and AI: Other Foundations of Human Computation & AICognitive Modeling & Cognitive Systems: Computational Creativity

Application Category

Economics, Online Markets and Human Computation: Trust and reliance of crowd workers and data experts on GenAISearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsSocial Networks and Social Media: Generative AI / large language models and their impact on social systems
📝 Abstract
Code comprehension is essential for brownfield programming tasks, in which developers maintain and enhance legacy code bases. Generative AI (GenAI) coding assistants such as GitHub Copilot have been shown to improve developer productivity, but their impact on code understanding is less clear. We replicate and extend a previous study by exploring both performance and comprehension in GenAI-assisted brownfield programming tasks. In a within-subjects experimental study, 18 computer science graduate students completed feature implementation tasks with and without Copilot. Results show that Copilot significantly reduced task time and increased the number of test cases passed. However, comprehension scores did not differ across conditions, revealing a comprehension-performance gap: participants passed more test cases with Copilot, but did not demonstrate greater understanding of the legacy codebase. Moreover, we failed to find a correlation between comprehension and task performance. These findings suggest that while GenAI tools can accelerate programming progress in a legacy codebase, such progress may come without an improved understanding of that codebase. We consider the implications of these findings for programming education and GenAI tool design.
Problem

Research questions and friction points this paper is trying to address.

Explores the comprehension-performance gap in GenAI-assisted legacy code programming
Investigates whether AI coding tools improve code understanding alongside productivity
Examines correlation between programming performance and code comprehension with AI assistance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Replicated study on GenAI-assisted legacy code tasks
Measured performance gains without comprehension improvement
Identified comprehension-performance gap in brownfield programming