Score
Designs, implements, debugs, and maintains software artifacts—applications, libraries, scripts, and services—by writing and integrating code and implementing algorithms and application logic. Performs hands‑on activities such as authoring code in relevant languages, writing tests, debugging, using version control and build tools, and setting up basic build and deployment workflows.
To address the challenges of simultaneously generating semantically consistent yet stylistically diverse multi-artifact programming exercises—namely source code, test specifications, and natural language descriptions—this paper proposes a compositional generation framework grounded in abstract syntax building blocks. The framework defines reusable syntactic abstractions and integrates templated mapping with multi-objective instantiation to ensure intent preservation and cross-modal co-generation. Its key innovations include: (i) enabling style-controllable, diverse outputs while guaranteeing semantic consistency; and (ii) providing a highly configurable generation interface that substantially reduces customization effort for new tasks. Experimental evaluation demonstrates that the approach outperforms existing baselines across three critical dimensions: generation quality, output diversity, and system extensibility.
This study addresses the limitation that analyzing prompts alone is insufficient for comprehensively evaluating developer interactions with AI programming agents. To overcome this, we propose a novel multidimensional interaction analysis framework termed "Say-Do-Understand," which integrates prompt data, screen activity, and comprehension metrics through a systematic five-stage end-to-end workflow. Employing an observational methodology, the analysis utilizes a prompt codebook, a screen activity coding scheme, and dual scoring rubrics. An empirical study involving ten experienced developers validates the proposed approach. Furthermore, four developer personas synthesizing task performance and comprehension levels are introduced to elucidate behavioral variations. Notably, the findings reveal that excessive reliance on agent self-checking significantly reduces developers' autonomous testing time.
This study addresses the limited understanding of how practitioners actually develop software engineering (SE) agents, particularly the lack of systematic investigation into the evolution of development workflows and core challenges. Through semi-structured interviews with 20 practitioners complemented by a survey of 80 respondents, this work proposes the first seven-stage workflow for SE agent development, revealing a paradigm shift toward “evaluation-driven iteration.” The research identifies that bottlenecks have moved beyond coding to non-coding tasks such as requirement specification, cross-role coordination, review, and deployment. It systematically characterizes six key challenges—including unreliable evaluation signals, accumulating comprehension debt, and behavioral drift induced by model updates—and synthesizes corresponding practical mitigation strategies.
研究通过分析GitHub Agentic Workflows的结构和维护方式,探讨了开发者如何定义和维护由AI代理执行的工作流程,并建议增加防御措施。
为解决自然语言工作流执行不可靠的问题,提出Artic编译器,将自然语言工作流转换为基于工件的工作流,提高任务解决率和一致性。
研究通过分析Claude Code插件市场中的1,926个仓库,探讨了AI编码代理插件的维护和共进化问题。
This study addresses the limited understanding of how Agent Control Files (ACFs)—instruction documents guiding autonomous coding agents—evolve, are maintained, and relate to code quality. Through large-scale repository mining, the authors reconstruct the commit-level evolutionary history of ACFs and propose, for the first time, a taxonomy of ACF changes grounded in software maintenance theory. By integrating qualitative content analysis, statistical testing, and code quality metrics, they empirically demonstrate how different types of maintenance activities differentially impact code quality and reveal dynamic patterns in these effects across the software development lifecycle. The findings provide both theoretical grounding and practical guidance for the governance of autonomous coding agents.
Bug fixing is a complex and time-consuming task in software development. Bug localization research tends to focus on the accuracy of automated tools that suggest source code files for developers to look at. However, little is known about how developers use these tools in practice. This paper reports on an ongoing qualitative user study. Eleven participants worked through four realistic bug localization tasks in a controlled environment and were given varying levels of support information offered by a specialized tool. Participants were asked to think aloud in a semi-structured interview session. The preliminary findings provide insight into three aspects of practice: how developers interact with tools, the role social and contextual information plays, and problem solving. The study demonstrates that bug localization is complex and suggests that the adoption of effective tools depends on more than their accuracy.
This study addresses the accountability deficit in agent development arising from the misalignment between platform controls and service provider terms. By analyzing four categories of tools and policy documents, we map workflow responsibilities and propose a novel grid model distinguishing verification mandates from executors. This framework reveals structural deficiencies in approval mechanisms, demonstrating that responsibility gaps have evolved from human oversight to inherent product attributes. Empirical findings indicate conflicting accountabilities across layers, contradictory attribution logic, and insufficient efficacy of approval artifacts. To support further research, we release a comprehensive dataset and validation scripts as open-source resources. Collectively, this work provides both theoretical grounding and empirical evidence necessary for reconstructing accountability frameworks in agent-based software systems, highlighting the urgent need to address systemic rather than incidental failures in current governance architectures.
This study addresses the disconnect between existing coding and computer-use agents, as well as the lack of visual interaction to assist software diagnosis and repair, by being the first to systematically investigate the role of visual feedback in this task. Methodologically, it integrates source-code-level execution, application screenshot analysis, and graphical interaction mechanisms to construct a benchmark environment spanning four domains, requiring agents to extract specification information from runtime interfaces and validate their modifications. The primary contribution lies in providing executable correctness evaluation criteria that systematically quantify the capability of state-of-the-art agents to accomplish software engineering tasks by combining code editing, command execution, and GUI-based visual feedback.