Score
Designs and implements automated pipelines, tooling, and configurations that build, test, package, and deploy software changes across environments (including artifact management, environment provisioning, and rollback mechanisms). Analyzes pipeline performance, reliability, and security to ensure repeatable, fast, and safe delivery of code.
This paper addresses the conceptual ambiguity, ill-defined boundaries, and lack of implementation standards between Infrastructure-as-Code (IaC) and Pipeline-as-Code in DevOps practice. To resolve these issues, we systematically delineate their respective roles and synergistic mechanisms within the DevOps ecosystem and propose a reusable, standardized IaC-driven CI/CD implementation framework. Our approach integrates Terraform for infrastructure provisioning, Ansible for configuration management, GitLab CI for pipeline orchestration, and Docker/Kubernetes for containerized deployment—enabling an end-to-end automated delivery pipeline. Empirical evaluation demonstrates 99.8% configuration change accuracy, reduces environment provisioning time from hours to minutes, and significantly improves deployment consistency and delivery efficiency.
To address the challenges of prolonged CI pipeline deployment cycles, error-prone manual configuration, and poor cross-project consistency, this paper proposes an automated pipeline configuration framework grounded in Infrastructure-as-Code (IaC) principles and templated configuration. The framework enables declarative definition and one-click generation of CI/CD pipelines via reusable YAML templates, a parameterized pipeline engine, and an integrated automation toolchain. Compared to conventional manual approaches, our method reduces average pipeline deployment time by 72% and decreases human configuration errors by 91%, while substantially improving consistency in build logic and execution environments across projects. Empirical validation across six open-source projects demonstrates the framework’s engineering practicality and methodological generality. It provides a reusable implementation model and actionable methodology for CI/CD automation, advancing scalable, maintainable, and reproducible software delivery practices.
This work addresses the growing complexity of CI/CD pipelines and the lack of structured analysis capabilities in existing tools for understanding their behavior, failures, and version evolution. The authors propose an innovative approach that uniquely integrates digital twin technology with BPMN-based modeling in DevOps contexts. By automatically parsing raw CI configurations and execution logs, the method constructs structured, high-level process models that enable pipeline visualization, failure traceability, and cross-version comparison. Evaluated across multiple open-source projects, the approach demonstrates effectiveness in monitoring, evolutionary analysis, and fault diagnosis, offering a modular and extensible foundational framework for the analysis and optimization of CI/CD pipelines.
Implementation discrepancies across software repository mining tools severely threaten the validity of empirical findings. Method: We conduct a dual-tool comparative analysis of 10 large-scale open-source projects, systematically identifying how minor implementation differences—such as commit parsing logic and author deduplication rules—induce up to 500% deviation in key metrics (e.g., commit count, developer count). We propose a “tool-level configuration + post-hoc normalization” co-optimization framework to mitigate metric divergence and perform multi-tool experiments, quantitative consistency assessment, and code-level root-cause analysis. Contribution/Results: We identify six technical sources undermining data validity and establish the first validity assessment paradigm for Mining Software Projects Research (MSPR) explicitly addressing tool heterogeneity—thereby enabling rigorous, reproducible, and comparable empirical software engineering studies.
This study presents the first empirical investigation into the evolution of CI/CD configurations in machine learning (ML) projects. Addressing the lack of understanding regarding how CI/CD configurations co-evolve with ML components, the authors analyze 508 open-source ML projects, 343 manually annotated commits, and 15,634 automated CI/CD commits. They propose a novel 14-category taxonomy capturing synergistic changes between CI/CD and ML components, develop a dedicated clustering tool to identify recurrent evolutionary patterns, and establish an empirically grounded model linking developer experience to CI/CD configuration modification behavior. Results show that 61.8% of CI/CD-related commits involve build strategy modifications; common anti-patterns—including dependency hardcoding and missing test frameworks—are identified; and senior developers modify CI/CD configurations more frequently and effectively than juniors, confirming the critical role of experience in CI/CD maintenance.
This work proposes a novel “Environment-in-the-Loop” paradigm that systematically integrates automated environment construction with code migration, addressing the common oversight in existing approaches of neglecting dynamic interaction with the target runtime environment. By leveraging large language model (LLM) agents to drive the co-evolution of code and its execution environment, the method combines static and dynamic environment analysis, API adaptation, and dependency resolution to significantly enhance both the accuracy and efficiency of code migration. The study demonstrates that joint migration of environment and code is not only beneficial but essential, offering a new pathway toward fully automated software evolution.
This work addresses the challenge developers face in efficiently authoring CI/CD configurations due to limited DevOps expertise by proposing a large language model (LLM)-based, context-aware generation approach. The method leverages both natural language descriptions and repository structure to automatically produce accurate and executable pipeline configurations for platforms such as GitHub Actions and GitLab CI/CD. Integrated with automated validation and human-in-the-loop feedback mechanisms, this framework is the first to combine repository context understanding with natural language-driven configuration synthesis. Experimental results demonstrate that the approach significantly lowers the barrier to DevOps adoption, markedly improves the accuracy and validity of generated configurations, and substantially reduces manual configuration effort.
This work addresses the unreliability of developer productivity dashboards, which often stems from ad hoc scripts that introduce undetected silent data gaps, eroding organizational trust. To resolve this, we propose a robust ELT pipeline grounded in DAG-based orchestration and the Medallion architecture, decoupling data extraction from transformation to preserve the immutability of raw data. Our approach introduces a state-driven dependency scheduling mechanism and, for the first time, treats metric pipelines as production-grade distributed systems. We emphasize the critical role of immutable raw history in enabling reliable metric redefinition. This methodology significantly enhances data reliability and freshness while effectively eliminating silent failures, thereby restoring organizational confidence in DevOps metrics.
Software maintenance remains heavily reliant on manual effort, resulting in high costs, low efficiency, and susceptibility to errors. This work proposes the first systematic research framework for transfer-based software maintenance, drawing inspiration from transfer learning. The framework establishes a comprehensive lifecycle model encompassing task identification, source system selection, cross-system data matching and adaptation, and validation of transferred outcomes. It explicitly delineates the core objectives and key challenges at each stage, integrating techniques from software engineering such as knowledge transfer, cross-project data alignment, and context-aware adaptation. By doing so, the framework introduces a novel paradigm for automating software maintenance and lays a solid theoretical foundation for the future development of supporting tools and methodologies.