Score
Designs, implements, tests, and maintains reliable, maintainable software by applying core practices such as modular design, algorithms and data structures, complexity analysis, version control, automated testing, debugging, and basic software architecture. Builds and evaluates artifacts like unit and integration tests, continuous integration/deployment pipelines, documentation, and code reviews to satisfy functional and non‑functional requirements (performance, scalability, reliability, maintainability, and basic security).
This work addresses the critical limitation of existing code generation systems—neglect of maintainability and poor adaptability to dynamic requirement changes—by pioneering maintainability as the primary optimization objective. We propose a novel code generation framework designed for continuous evolution, emphasizing high cohesion, low coupling, and easy adaptability. Methodologically, we (1) introduce MaintainBench, the first dynamic maintainability evaluation benchmark; (2) integrate waterfall-style phased governance, design-pattern-driven architectural generation, and multi-agent collaborative reasoning; and (3) incorporate a quantitative dynamic maintenance cost assessment model. Experimental results demonstrate a 14–30% improvement in maintainability metrics on MaintainBench, while simultaneously achieving superior pass@k functional correctness over baseline methods. All code and the MaintainBench benchmark are publicly released.
This work addresses the challenge of balancing software quality, testability, and maintainability under rapid iteration and frequent requirement changes. It proposes Algorithm-Driven Development (ADD), a novel approach that unifies requirements specification and technical design by using algorithm flowcharts as a single, coherent artifact. This integration enables end-to-end modeling of requirements, architecture, and testing. Leveraging this model, the system automatically generates high-coverage acceptance tests and incorporates continuous integration with code coverage feedback. Industrial adoption at Dassault Systèmes demonstrates that ADD achieves over 95% code coverage, substantially reduces defect density, and ensures a stable delivery cadence, outperforming conventional test-driven development and test-after approaches.
Software maintainability is frequently overlooked in requirements engineering, often addressed only implicitly through informal specifications or tool-based suggestions, without explicit goals or proactive management. This paper proposes a systematic framework for defining explicit maintainability requirements goals. It introduces the first adaptation of the QUPER model to the maintainability domain, integrating quantitative maintainability measurement tools and industry benchmarks to enable organizations to specify measurable, traceable, and actionable goals. The framework is developed and empirically validated using design science research methodology, with industrial case studies confirming its effectiveness in elevating maintainability’s priority within development decision-making. Key contributions include: (1) establishing maintainability as an explicit, goal-oriented requirement engineering concern; (2) providing the first QUPER-based approach for modeling and calibrating maintainability goals; and (3) delivering a practical, process-integrated solution deployable within real-world requirements engineering workflows.
Contract-based design (CbD) lacks systematic empirical evidence supporting its application in trustworthy software systems. Method: This study conducts the first systematic mapping study (SMS) of CbD across the full lifecycle and multiple dimensions for trustworthy systems, employing tri-database collaborative retrieval, collaborative review, and voting-based screening to analyze 288 primary studies via thematic coding and evidence aggregation. Contribution/Results: The study clarifies the breadth and depth of CbD adoption across domains, quantitatively assesses domain maturity distributions, and identifies critical gaps among theoretical modeling, automated verification, and industrial practice. It proposes six empirically grounded, verifiable research directions and delivers the first evidence-based CbD methodology framework and practical guidelines for trustworthy software engineering.
This paper investigates validity threats arising from toolchain selection in quantitative empirical software engineering. We formally replicate three high-impact studies by extracting identical project data using four widely adopted mining tools—Git, JIRA, GitHub API, and BIC—and conduct both quantitative and qualitative comparative analyses. Results demonstrate that subtle technical discrepancies across tools—including data modeling assumptions, event definitions, and temporal window handling—propagate and significantly undermine consistency in baseline datasets, statistical outcomes, and ultimately research conclusions. To our knowledge, this is the first systematic study to reveal the critical impact of tool choice on the robustness of empirical findings. We propose a practical framework comprising enhanced tool reusability, improved analytical transparency, and mandatory cross-tool validation. This work advances methodological rigor in software evolution research by highlighting and mitigating tool-induced validity threats.
This study addresses the limitation of existing templates in specification-driven development, which fail to evaluate specification clarity and completeness. To overcome this, we propose EPIC, a framework grounded in the ISO/IEC/IEEE 29148 standard that conducts quantitative assessments of open-source repositories. By distilling an optimal specification taxonomy encompassing ten quality dimensions and forty practices, EPIC guides developers in clarifying expectations and bridging specification gaps. Empirical evaluations demonstrate that high-quality specifications reduce the proportion of bug-fixing commits to 11.8% and yield a fourfold increase in the median number of contributors. These findings indicate that adopting rigorous specification practices significantly enhances both collaborative efficiency and software quality in open-source projects.
This study addresses the persistent occurrence of software defects after release, particularly in C/C++ and Java systems, whose underlying causes remain poorly understood. Through a large-scale empirical analysis of over 14,000 open-source projects, the work systematically compares pre-release and post-release defect characteristics using multidimensional metrics—including code complexity, size, change frequency, and development history—and employs statistical modeling to uncover key patterns. It reveals for the first time that post-release defects are significantly concentrated in legacy modules that undergo frequent modifications, with their root causes primarily stemming from dynamic evolutionary pressures rather than static code structure. Furthermore, such defects exhibit longer repair cycles and higher complexity, offering empirical grounding for targeted testing strategies and improved reliability assurance.
论文探讨了在软件代理时代,如何通过可信赖变更和责任拓扑结构来管理可扩展执行,以使组织能够接受并维持其变化。
本文提出了一种基于仓库的实现方法,通过自动接口更新和一致性检查减少有人和无人飞行器软件开发中跨域不一致问题。
本文探讨了大型语言模型在软件工程中基于测试的方法,通过分析87篇研究文献,区分并比较了不同测试驱动任务的特点和机制,提出了未来研究方向。