Score
Designs, implements, and evaluates development processes, toolchains, and code artifacts that enforce practices such as version control, automated testing, continuous integration/delivery, code review, modular architecture, release management, and documentation to produce maintainable, reliable, and scalable software. Analyzes codebases and workflows to identify gaps and implement improvements to quality, performance, and developer productivity.
Rapid delivery cycles exacerbate technical debt, undermining software maintainability. Method: This paper conducts a systematic case study of Meta to investigate sustainable code quality improvement mechanisms under high-frequency release regimes. We propose a layered code improvement framework—comprising spontaneous, institutionalized, and targeted refactoring—and develop a maintenance-priority–driven quality measurement model. We further introduce the novel “Improvement Contribution Badge” incentive mechanism. Leveraging source-code change history mining, industry case replication, and mixed qualitative-quantitative analysis, we identify that over 14% of code changes explicitly target maintainability enhancement. Results: Empirical evaluation demonstrates significant reductions in code complexity and measurable gains in developer productivity. Crucially, this work provides the first empirical validation that sustained, structured code improvement can concurrently preserve release velocity and ensure long-term maintainability—yielding a reusable methodology and a quantifiable governance framework for engineering practice.
This work addresses the challenge of balancing software quality, testability, and maintainability under rapid iteration and frequent requirement changes. It proposes Algorithm-Driven Development (ADD), a novel approach that unifies requirements specification and technical design by using algorithm flowcharts as a single, coherent artifact. This integration enables end-to-end modeling of requirements, architecture, and testing. Leveraging this model, the system automatically generates high-coverage acceptance tests and incorporates continuous integration with code coverage feedback. Industrial adoption at Dassault Systèmes demonstrates that ADD achieves over 95% code coverage, substantially reduces defect density, and ensures a stable delivery cadence, outperforming conventional test-driven development and test-after approaches.
Continuous Integration (CI) practices suffer from severe monitoring deficiencies: developers largely neglect critical metrics such as “build health” and “time-to-fix failed builds,” while mainstream CI services offer only weak native monitoring capabilities, forcing reliance on fragmented and often redundant third-party tools. Method: We conducted a triangulated investigation—including documentation analysis, developer surveys, functional audits of CI platforms, and case studies of open-source projects—to systematically identify cognitive gaps and practical monitoring needs. Contribution/Results: Our study provides the first empirical evidence that although over 80% of developers track test coverage, only a minority monitor build health or timeliness; further, all major CI services lack built-in multidimensional monitoring support. These findings establish an evidence-based foundation for designing next-generation CI monitoring frameworks and prioritizing tooling enhancements.
This study addresses the limitation of existing templates in specification-driven development, which fail to evaluate specification clarity and completeness. To overcome this, we propose EPIC, a framework grounded in the ISO/IEC/IEEE 29148 standard that conducts quantitative assessments of open-source repositories. By distilling an optimal specification taxonomy encompassing ten quality dimensions and forty practices, EPIC guides developers in clarifying expectations and bridging specification gaps. Empirical evaluations demonstrate that high-quality specifications reduce the proportion of bug-fixing commits to 11.8% and yield a fourfold increase in the median number of contributors. These findings indicate that adopting rigorous specification practices significantly enhances both collaborative efficiency and software quality in open-source projects.
This work addresses the critical limitation of existing code generation systems—neglect of maintainability and poor adaptability to dynamic requirement changes—by pioneering maintainability as the primary optimization objective. We propose a novel code generation framework designed for continuous evolution, emphasizing high cohesion, low coupling, and easy adaptability. Methodologically, we (1) introduce MaintainBench, the first dynamic maintainability evaluation benchmark; (2) integrate waterfall-style phased governance, design-pattern-driven architectural generation, and multi-agent collaborative reasoning; and (3) incorporate a quantitative dynamic maintenance cost assessment model. Experimental results demonstrate a 14–30% improvement in maintainability metrics on MaintainBench, while simultaneously achieving superior pass@k functional correctness over baseline methods. All code and the MaintainBench benchmark are publicly released.
This study addresses the limited understanding of how Agent Control Files (ACFs)—instruction documents guiding autonomous coding agents—evolve, are maintained, and relate to code quality. Through large-scale repository mining, the authors reconstruct the commit-level evolutionary history of ACFs and propose, for the first time, a taxonomy of ACF changes grounded in software maintenance theory. By integrating qualitative content analysis, statistical testing, and code quality metrics, they empirically demonstrate how different types of maintenance activities differentially impact code quality and reveal dynamic patterns in these effects across the software development lifecycle. The findings provide both theoretical grounding and practical guidance for the governance of autonomous coding agents.
This study addresses the lack of systematic understanding regarding how GitHub Actions workflows are used in real-world scenarios, how developers respond to workflow failures, and how these practices relate to project characteristics. Combining large-scale quantitative analysis of 258,300 workflow runs with qualitative case studies across 21 diverse repositories, this work identifies three typical patterns developers employ to handle workflow failures and uncovers a “configuration–usage gap”—where YAML configurations exist but workflows remain effectively unused. Furthermore, the study empirically validates five hypotheses linking project features to workflow usage intensity, revealing a significant positive correlation between high usage intensity and low failure rates. These findings provide actionable empirical evidence for improving CI/CD practices.
This study addresses the lack of systematic understanding regarding the evolution of GitHub Actions workflows. Through a mixed-methods approach, we conduct the first large-scale empirical analysis of over 3.4 million workflow file versions from more than 49,000 repositories spanning November 2019 to August 2025. We identify seven categories of conceptual changes and find that repositories typically contain a median of three workflow files, with 7.3% of workflows modified weekly—approximately 75% of which involve only a single change, predominantly in task configuration and specification. Our findings further indicate that current large language model (LLM) tools have not yet significantly influenced workflow maintenance frequency, offering empirical grounding for the design of fine-grained automated maintenance tools.
This work addresses the fragility of traditional requirements traceability in safety-critical software development, where reliance on external documentation often leads to silent breakdowns as code, requirements, and tests evolve independently. To overcome this, the paper proposes internalizing traceability as an intrinsic property of code structure by introducing language-native “Traceable” elements that enable compile-time verification of bidirectional links among requirements, implementations, and tests. The approach integrates code generation, metadata embedding, and build-time validation into the development workflow, providing proactive traceability assurance. When requirement changes disrupt traceability chains, the system automatically triggers warnings or build failures, thereby preventing traceability decay and significantly enhancing maintainability and reliability throughout software evolution.
本文针对Claude Code系统的安全和高效操作问题,通过提出四个核心原则和34章详细指南来解决。