Score
Designs, implements, and extends real‑time engine systems and runtime modules within and integrating with Unreal Engine, including C++ engine plugins, editor tools, rendering/physics subsystems, and Blueprint/C++ gameplay bindings. Builds and optimizes platform‑specific runtime behavior, networking and streaming systems, editor integration, and performance profiling/fixing for real‑time interactive applications.
Existing code-generating agents lack effective evaluation benchmarks within real-world, stateful, real-time C++ systems such as game engines. This work proposes the first evaluation platform built on Unreal Engine 5, comprising 110 C++ tasks extracted from nine real game repositories, spanning critical dimensions including game logic, networking, AI, and rendering. The framework employs behavior-driven testing and a multi-dimensional categorization scheme to automatically assess models via pass@1, measuring their ability to generate correct, compilable code within executable projects. Experimental results show that the best-performing model achieves a pass@1 rate of 55.5%, yet 31 tasks remain unsolved, highlighting significant challenges faced by current agents in deeply integrated development within complex C++ systems.
本文介绍了CraftBench-UE,一种在Unreal Engine中对编码代理进行确定性评估的工具,通过重建提交、执行检查来确保游戏功能正确实现,并比较了C++和Blueprint任务完成率。
VR development lacks quantitative, evidence-based criteria for selecting between Unreal Engine and Unity. Method: This study establishes a multidimensional empirical evaluation framework integrating rendering fidelity, computational efficiency, cross-platform compatibility, workflow productivity, and AI-enhanced capabilities (e.g., DLSS, LLM-assisted debugging), validated through systematic benchmarking and large-scale industrial case studies. Crucially, it pioneers the incorporation of AI-driven optimization techniques into engine performance attribution analysis and proposes a demand-aware, dynamic engine selection model grounded in project characteristics—such as immersion priority, hardware constraints, and team size. Contribution/Results: Findings indicate that high-fidelity VR applications favor Unreal Engine, whereas rapid-iteration or lightweight scenarios benefit more from Unity. The framework reduces selection-related trial-and-error costs by 37% and improves development efficiency by 2.1×, establishing a reusable, data-driven paradigm for engine selection in industrial VR deployment.
为了解决动作条件视频模型数据获取难题,本文提出基于虚幻引擎的两阶段合成数据生成管道,生成大规模动作条件多视角视频。
3D game development remains prohibitively high-barrier due to its reliance on programming, 3D modeling, and engine-specific configuration. Existing automated approaches are limited to 2D content, require manual integration, struggle with interactive logic and state management, and lack end-to-end support for mainstream engines (e.g., Unity, Unreal). Method: We propose the first zero-code, multi-agent collaborative framework for 3D game generation, leveraging multimodal large language models to orchestrate specialized agents for planning, code generation, automation, and debugging—enabling full pipeline translation from natural language specifications to executable C#-based Unity/Unreal projects. Contribution/Results: Evaluated on three prototype games, our framework reduces average development time by 91.4% and eliminates hand-written code entirely, marking a significant breakthrough in automating interactive 3D game creation.
针对游戏开发中视觉内容制作成本高、周期长的问题,Magpie系统通过分离游戏逻辑与图像生成,利用生成模型实现实时渲染,减少对完整视觉素材的依赖。
This study investigates the practical utility and capability boundaries of large language models (LLMs) in assisting with code refactoring and novel gameplay generation within real-world game development. Using a Python/Pygame-based endless runner game, GPT-4o was tasked with three localized refactoring operations and three cross-module gameplay generation challenges. Performance was evaluated through software metrics, unit tests, and manual playtesting. Results show that all refactoring tasks were correctly implemented, whereas only one of the three gameplay generation tasks was successfully integrated into the existing system. This work provides the first transparent case study demonstrating that LLMs excel at localized code modifications but face significant limitations when generating new features requiring coordination across multiple modules, offering empirical evidence and practical guidance for applying LLMs in game development contexts.
This study addresses the limitation of existing video benchmarks in evaluating models' capacity to follow fine-grained procedural world events. To this end, it constructs a novel benchmark grounded in replayable world records, generating videos through synchronized multi-view rendering and agent representations. Furthermore, the work proposes a vision-language model-based logic-rendering alignment metric that enables fine-grained consistency verification from procedural states to visual outputs. This approach effectively quantifies entity control, long-term memory, and interaction success rates, thereby establishing a rigorous evaluation standard for the visual fidelity of programmable world models.
This study addresses the difficulty of precisely editing existing executable worlds while preserving their original properties. We introduce the concept of "intervention depth" and propose the IGMWorld framework to enable hierarchical world editing and verification within game modding scenarios. Furthermore, we construct IGMBench, a benchmark comprising over one thousand criteria that systematically decouples generation, interaction, and editing capabilities. A multidimensional evaluation paradigm is designed, leveraging state-of-the-art coding agents integrated with deterministic executability, behavioral, and visual checks. Experimental results demonstrate that under optimal configurations, the task resolution rate reaches 78.2% with a 94.8% criterion pass rate, revealing the critical challenge of diminishing reliability in deep interventions.
This study addresses the lack of project-level code datasets and deterministic evaluation methodologies tailored to professional game engines, which has hindered the application of current AI techniques in full-scale game development. Leveraging open-source Game Jam projects and exploiting Godot’s text-based scene format and headless execution capabilities, the authors construct JamSet—the first project-level game code dataset comprising 8,133 validated projects—and JamBench, a benchmark suite of 300 projects. They further introduce a multidimensional evaluation framework featuring Structural Completeness Score (SCS) and Behavioral Alignment Score (BAS). Experiments reveal that state-of-the-art large language models achieve only a 5.7% execution pass rate on JamBench, highlighting architectural design as a critical bottleneck in generating complex game projects, while also demonstrating JamSet’s effectiveness for model training.