Score
Designs, implements, and evaluates software components and interfaces that connect game engines and their subsystems—especially graphics/rendering and physics pipelines—so assets, simulation state, and rendering passes interoperate correctly and perform well. Work includes developing engine plugins, renderer/physics bindings, pipeline and data-format adapters, editor/runtime glue, synchronization and update strategies, and platform-specific integration for engines such as Unity.
Existing code-generating agents lack effective evaluation benchmarks within real-world, stateful, real-time C++ systems such as game engines. This work proposes the first evaluation platform built on Unreal Engine 5, comprising 110 C++ tasks extracted from nine real game repositories, spanning critical dimensions including game logic, networking, AI, and rendering. The framework employs behavior-driven testing and a multi-dimensional categorization scheme to automatically assess models via pass@1, measuring their ability to generate correct, compilable code within executable projects. Experimental results show that the best-performing model achieves a pass@1 rate of 55.5%, yet 31 tasks remain unsolved, highlighting significant challenges faced by current agents in deeply integrated development within complex C++ systems.
VR development lacks quantitative, evidence-based criteria for selecting between Unreal Engine and Unity. Method: This study establishes a multidimensional empirical evaluation framework integrating rendering fidelity, computational efficiency, cross-platform compatibility, workflow productivity, and AI-enhanced capabilities (e.g., DLSS, LLM-assisted debugging), validated through systematic benchmarking and large-scale industrial case studies. Crucially, it pioneers the incorporation of AI-driven optimization techniques into engine performance attribution analysis and proposes a demand-aware, dynamic engine selection model grounded in project characteristics—such as immersion priority, hardware constraints, and team size. Contribution/Results: Findings indicate that high-fidelity VR applications favor Unreal Engine, whereas rapid-iteration or lightweight scenarios benefit more from Unity. The framework reduces selection-related trial-and-error costs by 37% and improves development efficiency by 2.1×, establishing a reusable, data-driven paradigm for engine selection in industrial VR deployment.
This study addresses the effective integration of generative tools into early-stage graphical asset creation in game development. Through a mixed-methods empirical study with 16 professional game designers and developers—including Likert-scale surveys, in-depth interviews, and statistical testing (p < .001)—we systematically identify key adoption requirements: strong user preference for tool intervention during conceptual design, demand for editable output formats and native IDE/engine integration, and a prevalent “quantity-first, refinement-later” workflow (mean quality tolerance: 0.17). Integration compatibility (mean rating: 3.5/5) and lack of universal data formats emerge as primary deployment bottlenecks. Based on these findings, we propose the first empirically grounded design guideline for generative graphics tools—specifying functional, interoperability, and workflow-aware principles. This work provides both theoretical foundations and actionable implementation pathways for the trustworthy integration of AIGC into industrial game development pipelines.
Physical engines (PEs) widely deployed in safety-critical systems—such as autonomous driving and medical robotics—frequently exhibit semantic-level physical failures, i.e., deviations from real-world physical behavior. However, existing testing approaches predominantly rely on white-box access and focus on crash detection, rendering them ineffective for identifying such subtle, semantics-driven failures. This paper presents the first large-scale empirical study to systematically characterize manifestations and root causes of physical failures in PEs, proposing the first fine-grained taxonomy. We comparatively evaluate diverse detection techniques and integrate deep learning, prompt engineering, and multimodal large language models to enable automated, semantic-level failure identification. We release PhysiXFails—an open-source benchmark dataset—and accompanying code, tools, and reproducible pipelines. Furthermore, informed by developer surveys, we propose actionable, deployable improvement strategies. Our work establishes a foundational framework—grounded in theory, empirical evidence, and practical implementation—for enhancing PE reliability.
This study addresses the lack of project-level code datasets and deterministic evaluation methodologies tailored to professional game engines, which has hindered the application of current AI techniques in full-scale game development. Leveraging open-source Game Jam projects and exploiting Godot’s text-based scene format and headless execution capabilities, the authors construct JamSet—the first project-level game code dataset comprising 8,133 validated projects—and JamBench, a benchmark suite of 300 projects. They further introduce a multidimensional evaluation framework featuring Structural Completeness Score (SCS) and Behavioral Alignment Score (BAS). Experiments reveal that state-of-the-art large language models achieve only a 5.7% execution pass rate on JamBench, highlighting architectural design as a critical bottleneck in generating complex game projects, while also demonstrating JamSet’s effectiveness for model training.
本文介绍了CraftBench-UE,一种在Unreal Engine中对编码代理进行确定性评估的工具,通过重建提交、执行检查来确保游戏功能正确实现,并比较了C++和Blueprint任务完成率。
为解决LLM编码代理在Unity项目中跨文件查询的问题,提出Unity Insight,一种代码-资产索引工具,显著减少查询时间和成本。
This study addresses the disconnect between existing coding and computer-use agents, as well as the lack of visual interaction to assist software diagnosis and repair, by being the first to systematically investigate the role of visual feedback in this task. Methodologically, it integrates source-code-level execution, application screenshot analysis, and graphical interaction mechanisms to construct a benchmark environment spanning four domains, requiring agents to extract specification information from runtime interfaces and validate their modifications. The primary contribution lies in providing executable correctness evaluation criteria that systematically quantify the capability of state-of-the-art agents to accomplish software engineering tasks by combining code editing, command execution, and GUI-based visual feedback.
This study addresses the inadequacy of existing benchmarks in evaluating coding agents’ ability to construct playable games. We propose SWE-Game, the first end-to-end development benchmark grounded in real executable games, comprising 41 Godot reference games and 247 tasks spanning the entire pipeline from brief generation to engine porting. The benchmark enables independent verification through shared instrumented interfaces and establishes a multidimensional evaluation framework integrating runtime checks, behavioral replay, and VLM-based visual scoring. Experimental results demonstrate that Opus5 achieves the best performance yet attains a construction score below 60. Runtime checks reach 92.59% accuracy, significantly outperforming video-based VLM judging, while visual scoring exhibits a correlation coefficient of 0.829 with human evaluations.
This study addresses the pervasive issues of structural disorder, poor maintainability, and frequent defects in game code, for which effective engineering solutions remain lacking. Employing a systematic literature review methodology, this work integrates code smell detection models, automated testing practices, and product line architecture techniques. It provides the first systematic mapping of the short-term development mindset and testing deficiencies unique to game development. By analyzing 34 relevant studies, this research identifies that existing detection and testing approaches have not been widely adopted, revealing a significant disconnect between academic findings and industrial practice along with substantial integration barriers. Ultimately, these insights establish a critical theoretical foundation for improving game software quality.