Score
Designs, implements, and evolves software applications and systems by writing code, defining architectures, creating tests and build/deployment automation, and producing supporting documentation. Also performs testing, debugging, maintenance, release management, deployment, and operational support to validate, launch, and sustain software products.
This study addresses the limited understanding of how practitioners actually develop software engineering (SE) agents, particularly the lack of systematic investigation into the evolution of development workflows and core challenges. Through semi-structured interviews with 20 practitioners complemented by a survey of 80 respondents, this work proposes the first seven-stage workflow for SE agent development, revealing a paradigm shift toward “evaluation-driven iteration.” The research identifies that bottlenecks have moved beyond coding to non-coding tasks such as requirement specification, cross-role coordination, review, and deployment. It systematically characterizes six key challenges—including unreliable evaluation signals, accumulating comprehension debt, and behavioral drift induced by model updates—and synthesizes corresponding practical mitigation strategies.
Software engineering lacks a formal theoretical foundation; existing research predominantly applies formal methods at the technical level rather than formally modeling software engineering itself. Method: This paper introduces the “meta-software engineering theory” paradigm—the first systematic effort to treat software engineering processes, entities (e.g., projects, modules, tests, milestones), and their interrelationships as formal objects. Leveraging object-oriented modeling integrated with formal methodologies, it constructs a structured theoretical prototype of core software engineering concepts, rigorously defining key abstractions and constraint relations. Contribution/Results: The work transcends conventional boundaries of formal method application, establishing a foundation for systematic, verifiable, and open-collaborative evolution of software engineering theory. It enables scalable development of a unified theoretical framework, supporting rigorous analysis, verification, and interoperable tooling across the software lifecycle.
Existing research lacks systematic methods to assess how requirements engineering (RE) impacts downstream development activities, hindering RE process optimization. Method: This paper proposes the first fitness-for-purpose RE impact assessment model, integrating a systematic literature review with multi-source empirical data to identify and structure 24 downstream development activities affected by requirements and 16 quantifiable attributes. Contribution/Results: The model bridges two critical gaps in requirements quality assessment—namely, the “activity dimension” and “measurability of impact”—by enabling empirical analysis of how specific requirements artifacts and processes concretely influence development practices. It provides a theoretically grounded framework and evidence-based decision support for precise, targeted optimization of the RE phase.
Scalability in assurance case (AC) development and maintenance for software product lines (SPLs) remains challenging due to the need to simultaneously accommodate variant diversity, perform evolution impact analysis, and enable certification evidence reuse. Method: This paper proposes a variant-aware formal approach that elevates AC construction to the product-line level. We define a variant-aware AC language and a template-based construction mechanism, enabling unified modeling of safety evidence and supporting property-level scalability and sustainable certification. Integrating variant logic, formal modeling, and model-driven engineering, we develop an automated toolchain for AC generation and maintenance. Contribution/Results: Empirical evaluation on a medical device SPL demonstrates that our approach significantly improves traceability accuracy and evidence reuse efficiency under evolutionary changes, thereby advancing scalable, maintainable, and certifiable SPL assurance.
This work addresses the challenge of balancing software quality, testability, and maintainability under rapid iteration and frequent requirement changes. It proposes Algorithm-Driven Development (ADD), a novel approach that unifies requirements specification and technical design by using algorithm flowcharts as a single, coherent artifact. This integration enables end-to-end modeling of requirements, architecture, and testing. Leveraging this model, the system automatically generates high-coverage acceptance tests and incorporates continuous integration with code coverage feedback. Industrial adoption at Dassault Systèmes demonstrates that ADD achieves over 95% code coverage, substantially reduces defect density, and ensures a stable delivery cadence, outperforming conventional test-driven development and test-after approaches.
本文提出了一种基于仓库的实现方法,通过自动接口更新和一致性检查减少有人和无人飞行器软件开发中跨域不一致问题。
This study addresses the limitation of existing templates in specification-driven development, which fail to evaluate specification clarity and completeness. To overcome this, we propose EPIC, a framework grounded in the ISO/IEC/IEEE 29148 standard that conducts quantitative assessments of open-source repositories. By distilling an optimal specification taxonomy encompassing ten quality dimensions and forty practices, EPIC guides developers in clarifying expectations and bridging specification gaps. Empirical evaluations demonstrate that high-quality specifications reduce the proportion of bug-fixing commits to 11.8% and yield a fourfold increase in the median number of contributors. These findings indicate that adopting rigorous specification practices significantly enhances both collaborative efficiency and software quality in open-source projects.
This study addresses the disconnect between existing coding and computer-use agents, as well as the lack of visual interaction to assist software diagnosis and repair, by being the first to systematically investigate the role of visual feedback in this task. Methodologically, it integrates source-code-level execution, application screenshot analysis, and graphical interaction mechanisms to construct a benchmark environment spanning four domains, requiring agents to extract specification information from runtime interfaces and validate their modifications. The primary contribution lies in providing executable correctness evaluation criteria that systematically quantify the capability of state-of-the-art agents to accomplish software engineering tasks by combining code editing, command execution, and GUI-based visual feedback.
This study addresses the challenges of requirement drift, perceptual deficits, and accountability ambiguity in coding agent iterations by proposing a Human-Agent-Virtual User engineering closed-loop framework. Leveraging multi-agent collaboration, version binding, and virtual user simulation, this approach enables end-to-end traceability and intent verification spanning from requirement confirmation to automated testing. The method establishes an auditable development lifecycle that effectively mitigates requirement drift while ensuring human oversight of final releases. Consequently, it achieves accountable agent-based application delivery and provides a reliable human-AI collaboration paradigm for complex software development.
本文探讨了大型语言模型在软件工程中基于测试的方法,通过分析87篇研究文献,区分并比较了不同测试驱动任务的特点和机制,提出了未来研究方向。