An Exploratory Case Study of LLM-Assisted Refactoring and Gameplay Feature Generation in an Endless Runner Game

📅 2026-06-19
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the practical utility and capability boundaries of large language models (LLMs) in assisting with code refactoring and novel gameplay generation within real-world game development. Using a Python/Pygame-based endless runner game, GPT-4o was tasked with three localized refactoring operations and three cross-module gameplay generation challenges. Performance was evaluated through software metrics, unit tests, and manual playtesting. Results show that all refactoring tasks were correctly implemented, whereas only one of the three gameplay generation tasks was successfully integrated into the existing system. This work provides the first transparent case study demonstrating that LLMs excel at localized code modifications but face significant limitations when generating new features requiring coordination across multiple modules, offering empirical evidence and practical guidance for applying LLMs in game development contexts.
📝 Abstract
Large language models (LLMs) are increasingly used to support software development, but their practical usefulness in applied game-development settings remains underexplored, especially when generated code must be integrated into an existing game software system. This paper presents an exploratory empirical case study of GPT-4o in a custom Python/Pygame endless runner. The study examines six selected development tasks: three localized refactoring tasks and three tasks involving gameplay feature generation. The resulting implementations were evaluated using software metrics, unit tests, and manual gameplay assessments. In this case study, all three selected refactoring tasks were completed successfully in functional terms, whereas only one of the three selected gameplay feature generation tasks resulted in a correctly integrated feature. The findings suggest that, in this setting, GPT-4o handled localized transformations more reliably than tasks requiring new gameplay interactions across multiple existing systems. Given the exploratory single-case design, these results are best interpreted as indicative observations rather than as generalizable evidence of category-level model performance. Overall, the paper contributes a transparent case-based account of the opportunities and limitations of LLM-assisted refactoring and gameplay feature generation in an existing game software system.
Problem

Research questions and friction points this paper is trying to address.

large language models
game development
code refactoring
feature generation
software integration
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM-assisted refactoring
gameplay feature generation
empirical case study
GPT-4o
software integration
🔎 Similar Papers
2024-07-24IEEE Transactions on GamesCitations: 0
J
Jan Wunderlich
IU International University of Applied Sciences, Erfurt, Germany
M
Markus Kleffmann
IU International University of Applied Sciences, Erfurt, Germany
S
Sebastian Lempert
IU International University of Applied Sciences, Erfurt, Germany