Score
Designs, implements, and operates the processes, tools, and artifacts that keep a system functioning correctly and securely over its lifecycle, including patching, bug fixes, refactoring, configuration updates, and dependency management. Analyzes logs, metrics, and failure modes to prioritize and plan corrective and preventive work, and builds deployment, rollback, and monitoring procedures to ensure continued availability and reliability.
Industrial applications heavily rely on open-source libraries, yet stalled community maintenance frequently leaves vulnerabilities unpatched for extended periods, posing critical software supply chain security risks. Existing approaches suffer from label scarcity, sparse feature representations, and incomplete modeling of transitive dependency relationships, hindering practical deployment in industrial settings. This paper proposes the first maintenance-activity monitoring framework that jointly models direct and transitive dependencies. It constructs fine-grained maintenance metrics from multi-source repository metadata—including commits, releases, issues, and pull requests—and introduces a graph propagation model to quantify the cross-dependency transmission of maintenance decay. Crucially, the method operates without manual labeling. Evaluated across multiple enterprise projects, it achieves early warning of high-risk stagnant libraries 3–6 months in advance, substantially reducing manual auditing effort and significantly enhancing the security and maintainability of open-source dependency ecosystems.
To address the challenges of standardizing Site Reliability Engineering (SRE) practices in heterogeneous environments and balancing system reliability with development agility, this paper proposes a customizable SRE process framework. The framework integrates automated operations, multidimensional observability (metrics, logs, traces), error-budget-driven governance, standardized incident response, and progressive delivery (canary and blue-green deployments). It is designed for cross-technology-stack adaptability, enabling contextual implementation of core SRE principles. Evaluated in production systems, the framework reduced mean time to recovery by 42%, decreased unplanned outages by 67%, lowered operational staffing requirements by 35%, and achieved 99.99% service availability. Its primary contribution is the first methodology for customizing SRE processes specifically for heterogeneous environments, empirically demonstrating synergistic improvements in both system reliability and operational efficiency.
Prior work lacks empirical characterization of problem-solving processes in software development. Method: Integrating grounded theory coding, sequential pattern mining, and multidimensional statistical analysis on 356 Mozilla Firefox issue reports, this study extracts fine-grained, reusable problem-solving process patterns from collaborative textual artifacts. Contribution/Results: We identify 47 empirically grounded process patterns—challenging the traditional linear assumption by revealing pervasive nonlinearity: 73% of fixes involve iterative backtracking or parallel activities. The resulting process landscape and pattern catalog systematically characterize distributional regularities across issue types, defect categories, and repair durations. This advances understanding of real-world engineering complexity and provides an evidence-based foundation for process optimization, collaborative tool design, and developer support.
Continuous Integration (CI) practices suffer from severe monitoring deficiencies: developers largely neglect critical metrics such as “build health” and “time-to-fix failed builds,” while mainstream CI services offer only weak native monitoring capabilities, forcing reliance on fragmented and often redundant third-party tools. Method: We conducted a triangulated investigation—including documentation analysis, developer surveys, functional audits of CI platforms, and case studies of open-source projects—to systematically identify cognitive gaps and practical monitoring needs. Contribution/Results: Our study provides the first empirical evidence that although over 80% of developers track test coverage, only a minority monitor build health or timeliness; further, all major CI services lack built-in multidimensional monitoring support. These findings establish an evidence-based foundation for designing next-generation CI monitoring frameworks and prioritizing tooling enhancements.
This work addresses the limitations of existing large language models, which are typically confined to isolated tasks and struggle to integrate into industrial-scale, multi-stage security workflows. To bridge this gap, the authors propose the first role-based multi-agent framework tailored to the entire vulnerability lifecycle, incorporating specialized agents—Planner, Analyzer, Fixer, and Verifier—augmented with CodeQL static analysis for enhanced precision. By introducing a role-oriented multi-agent architecture into end-to-end vulnerability management, this approach effectively aligns the capabilities of large models with real-world security engineering demands. Evaluated on 25 real-world C/C++ vulnerabilities, the system achieves a detection accuracy of 44%—comparable to GPT-5.5—and a repair accuracy of 19%, offering a practical and collaborative paradigm for intelligent security operations.
This study addresses the challenge of effectively monitoring early-stage agent systems, where structural flaws often obscure task-level errors. The authors propose a three-dimensional (quality, suitability, efficiency) and three-granularity (intra-run, inter-run, structural) monitoring and triaging framework tailored for low-maturity agent systems. They introduce a novel system maturity staging model based on the coefficient of variation and monitoring granularity, integrated with a severity classification adapted from FMEA to guide human review. The resulting transferable monitoring architecture supports document-driven, multi-stage workflows, enhanced by a synthetic testbed with controlled error injection. Experimental results demonstrate that structural defects significantly mask task-level signals; 97% of issues can be automatically traced, with only 2% requiring human intervention, and each granularity level precisely identifies its corresponding defect type (coefficients of variation: 0.02, 1.25, and 0.00, respectively).
This study addresses the lack of systematic understanding regarding how GitHub Actions workflows are used in real-world scenarios, how developers respond to workflow failures, and how these practices relate to project characteristics. Combining large-scale quantitative analysis of 258,300 workflow runs with qualitative case studies across 21 diverse repositories, this work identifies three typical patterns developers employ to handle workflow failures and uncovers a “configuration–usage gap”—where YAML configurations exist but workflows remain effectively unused. Furthermore, the study empirically validates five hypotheses linking project features to workflow usage intensity, revealing a significant positive correlation between high usage intensity and low failure rates. These findings provide actionable empirical evidence for improving CI/CD practices.
Bug fixing is a complex and time-consuming task in software development. Bug localization research tends to focus on the accuracy of automated tools that suggest source code files for developers to look at. However, little is known about how developers use these tools in practice. This paper reports on an ongoing qualitative user study. Eleven participants worked through four realistic bug localization tasks in a controlled environment and were given varying levels of support information offered by a specialized tool. Participants were asked to think aloud in a semi-structured interview session. The preliminary findings provide insight into three aspects of practice: how developers interact with tools, the role social and contextual information plays, and problem solving. The study demonstrates that bug localization is complex and suggests that the adoption of effective tools depends on more than their accuracy.