Score
Designs, implements, tests, debugs, and maintains software by writing source code, scripts, libraries, tools, and APIs using programming languages and development tooling. Produces correct, efficient, readable, and maintainable code through modular design, automated tests, debugging, profiling, and refactoring.
To address the challenges of simultaneously generating semantically consistent yet stylistically diverse multi-artifact programming exercises—namely source code, test specifications, and natural language descriptions—this paper proposes a compositional generation framework grounded in abstract syntax building blocks. The framework defines reusable syntactic abstractions and integrates templated mapping with multi-objective instantiation to ensure intent preservation and cross-modal co-generation. Its key innovations include: (i) enabling style-controllable, diverse outputs while guaranteeing semantic consistency; and (ii) providing a highly configurable generation interface that substantially reduces customization effort for new tasks. Experimental evaluation demonstrates that the approach outperforms existing baselines across three critical dimensions: generation quality, output diversity, and system extensibility.
This work addresses the critical limitation of existing code generation systems—neglect of maintainability and poor adaptability to dynamic requirement changes—by pioneering maintainability as the primary optimization objective. We propose a novel code generation framework designed for continuous evolution, emphasizing high cohesion, low coupling, and easy adaptability. Methodologically, we (1) introduce MaintainBench, the first dynamic maintainability evaluation benchmark; (2) integrate waterfall-style phased governance, design-pattern-driven architectural generation, and multi-agent collaborative reasoning; and (3) incorporate a quantitative dynamic maintenance cost assessment model. Experimental results demonstrate a 14–30% improvement in maintainability metrics on MaintainBench, while simultaneously achieving superior pass@k functional correctness over baseline methods. All code and the MaintainBench benchmark are publicly released.
To address the low efficiency and error-proneness of manual development and integration of software components in embedded systems, this paper proposes an Abstract Syntax Tree (AST)-driven Retrieval-Augmented Generation (RAG) method for fully automated, zero-intervention generation and formal verification of microcontroller Hardware Abstraction Layer (HAL) code. Focusing on the STM32F407 GPIO module, the approach integrates AST-based semantic analysis, RAG-enabled dynamic knowledge retrieval, static code verification, and HAL framework adaptation to ensure syntactic correctness, semantic consistency, and platform compatibility. Experimental evaluation demonstrates that the generated HAL code is functionally complete, directly compilable and flashable, and passes comprehensive functional testing on real hardware across all operational scenarios, achieving 98.7% accuracy. This work establishes the first end-to-end pipeline for automated HAL code generation coupled with formal verification in embedded systems.
本文提出了一种基于Eclipse JDT API的自动重构工具,通过引入辅助布尔变量转换含有break和continue语句的Java代码结构,以改善代码结构并支持进一步的自动化重构。
Coverage-driven testing often fails to capture semantic behavior and misses deep-seated defects. To address this, this paper proposes a behavior-oriented test generation method leveraging code comments (e.g., Javadoc). First, functional specifications are extracted from natural-language comments and formally modeled as executable test objectives. Second, search-based test generation is integrated with context-aware assertion synthesis to automatically generate named, behaviorally meaningful test cases. This work is the first to explicitly transform code comments into executable test goals, thereby overcoming the semantic blindness of conventional coverage metrics and enabling precise, behavior-level testing. Evaluated on a benchmark of 118 Java classes, our approach significantly improves behavioral coverage and successfully detects multiple previously unknown defects that evade state-of-the-art coverage-based tools.
本文探讨了大型语言模型在软件工程中基于测试的方法,通过分析87篇研究文献,区分并比较了不同测试驱动任务的特点和机制,提出了未来研究方向。
研究通过分析Claude Code插件市场中的1,926个仓库,探讨了AI编码代理插件的维护和共进化问题。
研究探讨了代码可维护性对LLM生成单元测试有效性的影响,使用CodeHealth指标评估,并发现其与输入token数负相关。
This study addresses growing industry concerns about the practicality and naturalness of code generated by large language models (LLMs) by systematically examining the usage patterns and defect associations of LLM-generated code and comments in active enterprise and community repositories from 2021 to 2025. For the first time, it contrasts the distribution of LLM-generated content between these two repository types through an empirical analysis integrating multiple detection tools, code clone detection, syntactic quality assessment, and manually labeled defect data. The findings reveal that the proportion of LLM-generated code has steadily declined over time and is predominantly confined to test cases, while comment generation remains stable yet exhibits low syntactic correctness. Enterprise repositories incorporate more LLM-generated content overall, which shows virtually no direct association with known defects, suggesting that such content is characterized by low risk but high functional limitations in real-world practice.
Traditional testing approaches struggle to uncover deep logical flaws that violate explicit or implicit constraints, particularly when discrepancies exist between requirements and implementation. This work proposes a novel method that leverages large language models to automatically generate Alloy formal specifications from both requirement documents and source code, using these specifications as an intermediate representation to derive executable test cases. By integrating large language models, formal methods, static analysis, and automated testing, the approach effectively exposes constraint-level defects missed by conventional test generation techniques. Empirical evaluation reveals a real-world vulnerability in the Flipper library and uncovers undocumented implicit abstractions in Cerberus. Moreover, tests derived from code-generated specifications demonstrate superior stability compared to those based solely on requirements.