Score
Designs, implements, and maintains declarative, version-controlled infrastructure definitions and automated provisioning/deployment pipelines using infrastructure‑as‑code tools (e.g., Terraform, Pulumi), and builds the related CI/CD, configuration and governance artifacts that manage environments. Also designs and deploys automated model training, serving and inference workflows and integrates AI‑assisted or autonomous development tooling to generate, test, and govern code and deployment processes.
This paper addresses the conceptual ambiguity, ill-defined boundaries, and lack of implementation standards between Infrastructure-as-Code (IaC) and Pipeline-as-Code in DevOps practice. To resolve these issues, we systematically delineate their respective roles and synergistic mechanisms within the DevOps ecosystem and propose a reusable, standardized IaC-driven CI/CD implementation framework. Our approach integrates Terraform for infrastructure provisioning, Ansible for configuration management, GitLab CI for pipeline orchestration, and Docker/Kubernetes for containerized deployment—enabling an end-to-end automated delivery pipeline. Empirical evaluation demonstrates 99.8% configuration change accuracy, reduces environment provisioning time from hours to minutes, and significantly improves deployment consistency and delivery efficiency.
This study addresses the high manual overhead faced by DevOps teams in managing multi-interface cloud infrastructures. We propose and systematically evaluate an LLM-driven AI agent framework for automation. Methodologically, the agent unifies heterogeneous interfaces—including SDKs, CLIs, Infrastructure-as-Code (IaC) tools, and web portals—to support core tasks such as configuration deployment, monitoring/alerting, and incident remediation. Key contributions include: (1) the first evaluation framework specifically designed for AI agents in cloud infrastructure management; (2) identification and systematic mitigation of three critical bottlenecks—interface semantic gaps, action execution reliability, and security constraint compliance; and (3) domain-specific optimization strategies validated in real-world deployments, demonstrating both task feasibility and cross-scenario generalizability. Our work establishes a reusable methodology and empirical benchmark for AI-native cloud operations.
In modern CI/CD pipelines, manual intervention in unstable test diagnosis, rollback decisions, feature flag tuning, and canary promotion introduces release delays and operational overhead. To address this, we propose an AI-augmented autonomous software delivery framework that integrates large language models (LLMs) with policy-constrained autonomous agents, yielding a reference architecture for agent-based decision-making bounded by formal policies. Our approach introduces: (1) a taxonomy of deployment decisions; (2) policy-as-code guardrails enforcing safety and compliance; (3) a tiered trust framework governing agent autonomy; and (4) a DORA-metrics-driven, verifiable evaluation methodology. Evaluated in a React 19 microservices environment, the framework significantly reduces deployment latency and manual intervention frequency, improves release velocity and system reliability, and ensures auditable, formally verifiable autonomous decision paths.
To address the challenges of prolonged CI pipeline deployment cycles, error-prone manual configuration, and poor cross-project consistency, this paper proposes an automated pipeline configuration framework grounded in Infrastructure-as-Code (IaC) principles and templated configuration. The framework enables declarative definition and one-click generation of CI/CD pipelines via reusable YAML templates, a parameterized pipeline engine, and an integrated automation toolchain. Compared to conventional manual approaches, our method reduces average pipeline deployment time by 72% and decreases human configuration errors by 91%, while substantially improving consistency in build logic and execution environments across projects. Empirical validation across six open-source projects demonstrates the framework’s engineering practicality and methodological generality. It provides a reusable implementation model and actionable methodology for CI/CD automation, advancing scalable, maintainable, and reproducible software delivery practices.
This work addresses the challenge developers face in efficiently authoring CI/CD configurations due to limited DevOps expertise by proposing a large language model (LLM)-based, context-aware generation approach. The method leverages both natural language descriptions and repository structure to automatically produce accurate and executable pipeline configurations for platforms such as GitHub Actions and GitLab CI/CD. Integrated with automated validation and human-in-the-loop feedback mechanisms, this framework is the first to combine repository context understanding with natural language-driven configuration synthesis. Experimental results demonstrate that the approach significantly lowers the barrier to DevOps adoption, markedly improves the accuracy and validity of generated configurations, and substantially reduces manual configuration effort.
Existing Infrastructure-as-Code (IaC) repair approaches often rely on manual intervention or are prone to hallucination, compromising repair validity. This work proposes TerraRepair, the first large language model agent for IaC repair that incorporates a tool-anchoring mechanism. TerraRepair retrieves Terraform dependency context, queries provider schemas, and re-invokes Checkov and Trivy post-repair to ensure correctness. Crucially, when essential contextual information is missing, it proactively escalates the issue rather than generating speculative fixes. Experimental results on an AWS benchmark demonstrate that TerraRepair significantly improves verified repair rates—increasing them from 26.6% to 78.4% for Checkov and from 44.8% to 72.4% for Trivy—with the majority of repairs validated as correct through human evaluation.
This study addresses the limited understanding of how Agent Control Files (ACFs)—instruction documents guiding autonomous coding agents—evolve, are maintained, and relate to code quality. Through large-scale repository mining, the authors reconstruct the commit-level evolutionary history of ACFs and propose, for the first time, a taxonomy of ACF changes grounded in software maintenance theory. By integrating qualitative content analysis, statistical testing, and code quality metrics, they empirically demonstrate how different types of maintenance activities differentially impact code quality and reveal dynamic patterns in these effects across the software development lifecycle. The findings provide both theoretical grounding and practical guidance for the governance of autonomous coding agents.
This work addresses critical challenges in configuration management for large language model (LLM)-based coding agents, including configuration reuse ambiguities, unclear permission boundaries, and inadequate versioning. To tackle these issues, the authors propose Rel(AI)Build—the first deterministic, tool-agnostic configuration governance framework specifically designed for LLM coding agents. Treating agent definitions as managed supply chains, Rel(AI)Build enforces configuration integrity through SHA-256 content addressing, HMAC-signed lockfiles, hash-chain audit logs, hierarchical access controls, and a state-machine-driven development workflow. The framework also supports multi-IDE target compilation. Empirical evaluation demonstrates that Rel(AI)Build effectively preserves configuration immutability under adversarial compliance tests, thereby validating its reliability and security guarantees.