Score
Designs and implements automated pipelines, test suites, and verification tooling to validate that releases, patches, and deployment artifacts are correctly built, installed, configured, and behave as expected before, during, and after launch. Builds and analyzes pre- and post-deployment checks—such as smoke and integration tests, configuration/environment verification, artifact/source integrity and patch verification, rollback checks, and observability/metric comparisons—to detect regressions or deployment failures and to drive gating or remediation.
This paper addresses the conceptual ambiguity, ill-defined boundaries, and lack of implementation standards between Infrastructure-as-Code (IaC) and Pipeline-as-Code in DevOps practice. To resolve these issues, we systematically delineate their respective roles and synergistic mechanisms within the DevOps ecosystem and propose a reusable, standardized IaC-driven CI/CD implementation framework. Our approach integrates Terraform for infrastructure provisioning, Ansible for configuration management, GitLab CI for pipeline orchestration, and Docker/Kubernetes for containerized deployment—enabling an end-to-end automated delivery pipeline. Empirical evaluation demonstrates 99.8% configuration change accuracy, reduces environment provisioning time from hours to minutes, and significantly improves deployment consistency and delivery efficiency.
This work addresses the challenge of silent updates to large language models (LLMs) by service providers, which often occur without version changes and can lead to behavioral drift and functional regressions, while existing mechanisms lack deployment-side control over compatibility governance. Framing LLM updates as a software supply chain governance problem, this study proposes a deployment-side control framework that defines rule-based production contracts, constructs risk-category-oriented test suites, and enforces compatibility gates to validate model safety and performance prior to updates. Experimental results demonstrate that the approach effectively uncovers fine-grained regressions missed by aggregate metrics, while also highlighting critical challenges in test design, threshold calibration, and drift attribution.
Continuous Integration (CI) practices suffer from severe monitoring deficiencies: developers largely neglect critical metrics such as “build health” and “time-to-fix failed builds,” while mainstream CI services offer only weak native monitoring capabilities, forcing reliance on fragmented and often redundant third-party tools. Method: We conducted a triangulated investigation—including documentation analysis, developer surveys, functional audits of CI platforms, and case studies of open-source projects—to systematically identify cognitive gaps and practical monitoring needs. Contribution/Results: Our study provides the first empirical evidence that although over 80% of developers track test coverage, only a minority monitor build health or timeliness; further, all major CI services lack built-in multidimensional monitoring support. These findings establish an evidence-based foundation for designing next-generation CI monitoring frameworks and prioritizing tooling enhancements.
This work addresses the growing complexity of CI/CD pipelines and the lack of structured analysis capabilities in existing tools for understanding their behavior, failures, and version evolution. The authors propose an innovative approach that uniquely integrates digital twin technology with BPMN-based modeling in DevOps contexts. By automatically parsing raw CI configurations and execution logs, the method constructs structured, high-level process models that enable pipeline visualization, failure traceability, and cross-version comparison. Evaluated across multiple open-source projects, the approach demonstrates effectiveness in monitoring, evolutionary analysis, and fault diagnosis, offering a modular and extensible foundational framework for the analysis and optimization of CI/CD pipelines.
Automated software environment deployment remains challenging due to complex dependencies, heterogeneous build systems, and insufficient documentation, often hindering reproducibility. This work proposes a three-tier pyramid model of environment maturity grounded in executable evidence, with successful execution of the main entry point as the highest validation criterion—surpassing the limitations of traditional weak-signal assessments. The approach integrates large language models with an execution-feedback loop, leveraging hierarchical validation, incremental repair, and deep understanding of project structure to iteratively construct runnable environments. Evaluated on four public benchmarks, the method substantially outperforms existing techniques, achieving up to a 79.6% improvement overall and a 66.7% gain on C/C++ projects, while successfully configuring 11 to 30 previously unsolvable environment instances for the first time.
This work addresses the lack of security and verifiability in large language models for project-level code generation by proposing and evaluating an end-to-end Detect–Repair–Verify (DRV) workflow tailored for multilingual web applications. The approach generates executable code at three granularities—project, requirement, and function—integrating static and dynamic analysis, automated repair, and test-driven verification. Under unified resource constraints, the study systematically compares generative, single-round, and iterative variants of DRV. It introduces the first project-level benchmark for secure code generation that supports multiple prompting granularities, enabling a comprehensive evaluation of DRV’s efficacy. The findings reveal limitations in using vulnerability reports to guide repairs and identify common post-repair failure modes such as regressions and semantic drift. Experimental results demonstrate that the iterative DRV variant significantly enhances security while preserving functional correctness.
This study addresses the persistent occurrence of software defects after release, particularly in C/C++ and Java systems, whose underlying causes remain poorly understood. Through a large-scale empirical analysis of over 14,000 open-source projects, the work systematically compares pre-release and post-release defect characteristics using multidimensional metrics—including code complexity, size, change frequency, and development history—and employs statistical modeling to uncover key patterns. It reveals for the first time that post-release defects are significantly concentrated in legacy modules that undergo frequent modifications, with their root causes primarily stemming from dynamic evolutionary pressures rather than static code structure. Furthermore, such defects exhibit longer repair cycles and higher complexity, offering empirical grounding for targeted testing strategies and improved reliability assurance.