Score
Designs, builds, and operates staged rollout and canary release processes that deploy new software versions to a subset of users or infrastructure, including orchestration and CI/CD pipelines that progressively shift traffic and automate testing. Develops canary deployment strategies, patterns, analysis and rollback mechanisms, and monitoring-driven decision logic to validate behavior, minimize user impact, and enable safe automated or manual rollbacks.
This work addresses the challenge of silent updates to large language models (LLMs) by service providers, which often occur without version changes and can lead to behavioral drift and functional regressions, while existing mechanisms lack deployment-side control over compatibility governance. Framing LLM updates as a software supply chain governance problem, this study proposes a deployment-side control framework that defines rule-based production contracts, constructs risk-category-oriented test suites, and enforces compatibility gates to validate model safety and performance prior to updates. Experimental results demonstrate that the approach effectively uncovers fine-grained regressions missed by aggregate metrics, while also highlighting critical challenges in test design, threshold calibration, and drift attribution.
This work addresses the fragility, inefficiency, and strong platform coupling commonly found in CI/CD pipelines for legacy COBOL systems, which often result in high maintenance costs and vendor lock-in. To overcome these challenges, the authors propose a portable CI/CD architecture tailored for highly secure and compliance-driven environments. The approach leverages OCI-compliant container images preloaded with COBOL toolchains, introduces a platform abstraction layer, integrates multiple repositories, and employs Groovy script refactoring to achieve platform-agnostic continuous integration and delivery. Empirical evaluation demonstrates that the proposed solution significantly enhances efficiency—reducing pipeline execution time by 82%—while simultaneously improving system portability, security, and maintainability. This architecture offers a reusable paradigm for modernizing legacy COBOL applications within regulated domains.
Traditional Jenkins controllers suffer from resource overloading and reduced reliability due to direct execution of build tasks. To address this, we propose a lightweight CI/CD architecture that containerizes the Jenkins controller and offloads all build execution to remote Docker hosts via secure SSH connections—effectively decoupling orchestration from build execution. The architecture incorporates atomic deployments, timestamped artifact backups, immutable artifact packaging, and automated notification mechanisms. Technically, it integrates persistent volumes, containerized build environments, and declarative pipelines. Experimental evaluation demonstrates a significant reduction in controller CPU and memory utilization, a 32% increase in build throughput, and a 41% decrease in artifact delivery latency. The solution delivers high stability, scalability, and low operational overhead, making it particularly suitable for small- to medium-scale DevOps environments.
In modern CI/CD pipelines, manual intervention in unstable test diagnosis, rollback decisions, feature flag tuning, and canary promotion introduces release delays and operational overhead. To address this, we propose an AI-augmented autonomous software delivery framework that integrates large language models (LLMs) with policy-constrained autonomous agents, yielding a reference architecture for agent-based decision-making bounded by formal policies. Our approach introduces: (1) a taxonomy of deployment decisions; (2) policy-as-code guardrails enforcing safety and compliance; (3) a tiered trust framework governing agent autonomy; and (4) a DORA-metrics-driven, verifiable evaluation methodology. Evaluated in a React 19 microservices environment, the framework significantly reduces deployment latency and manual intervention frequency, improves release velocity and system reliability, and ensures auditable, formally verifiable autonomous decision paths.
This study presents the first empirical investigation into the evolution of CI/CD configurations in machine learning (ML) projects. Addressing the lack of understanding regarding how CI/CD configurations co-evolve with ML components, the authors analyze 508 open-source ML projects, 343 manually annotated commits, and 15,634 automated CI/CD commits. They propose a novel 14-category taxonomy capturing synergistic changes between CI/CD and ML components, develop a dedicated clustering tool to identify recurrent evolutionary patterns, and establish an empirically grounded model linking developer experience to CI/CD configuration modification behavior. Results show that 61.8% of CI/CD-related commits involve build strategy modifications; common anti-patterns—including dependency hardcoding and missing test frameworks—are identified; and senior developers modify CI/CD configurations more frequently and effectively than juniors, confirming the critical role of experience in CI/CD maintenance.
This study addresses the lack of systematic understanding in the configuration and maintenance of CI/CD caching, which imposes a significant burden on developers despite its benefits for build efficiency. Through a large-scale empirical analysis of 952 repositories on GitHub Actions—encompassing 1,556 workflow files and over ten thousand cache-related changes—the authors employ code mining, configuration analysis, commit tracing, and statistical modeling to uncover real-world caching practices, evolutionary patterns, and human-bot collaboration in maintenance. The findings reveal that cache adopters are more active, caching strategies are diverse and frequently adjusted, and build- and test-related tasks evolve rapidly. Manual interventions primarily address misconfigurations, whereas version upgrades are predominantly automated by bots. The work quantifies the maintenance overhead of caching and provides empirical foundations for improving developer tooling.
This work addresses the inefficiency and error-proneness of manually aggregating change descriptions and impact scopes in cloud-native CI/CD pipelines during multi-task, multi-author collaborative releases. To tackle this challenge, the authors propose a novel approach that integrates semantic commit filtering, large language model (LLM)-driven structured summarization, and static task dependency analysis—marking the first integration of LLMs with pipeline dependency analysis to automatically generate stakeholder-oriented, categorized change reports. The system has been implemented within GitHub Actions and Tekton and deployed in a production environment comprising over 20 pipelines and 60 tasks. Empirical results demonstrate significant improvements in the accuracy and timeliness of release communication, outperforming existing tools such as SmartNote and VerLog.
This work addresses the challenge of repeated re-certification in traditional canary deployments for safety-critical embodied intelligent agents, which arises from changes to the system’s cryptographic identity. The authors propose ICAN-Deploy, a middleware that decouples immutable capability names—hashed to serve as stable identities—from mutable capability versions, thereby preserving identity hashes throughout the canary window and enabling, for the first time, identity-stable canary deployment. The approach integrates a state machine design, a runtime governance layer, AST-based static analysis, closed-form proofs, and TLA+ model checking to guarantee safety and correctness. Empirical validation on a Franka Panda robotic arm in MuJoCo across 100 deployments demonstrates zero identity drift, with entry latency falling within a 95% BCa confidence interval of [1.52, 2.01] milliseconds.
This work addresses the growing complexity of CI/CD pipelines and the lack of structured analysis capabilities in existing tools for understanding their behavior, failures, and version evolution. The authors propose an innovative approach that uniquely integrates digital twin technology with BPMN-based modeling in DevOps contexts. By automatically parsing raw CI configurations and execution logs, the method constructs structured, high-level process models that enable pipeline visualization, failure traceability, and cross-version comparison. Evaluated across multiple open-source projects, the approach demonstrates effectiveness in monitoring, evolutionary analysis, and fault diagnosis, offering a modular and extensible foundational framework for the analysis and optimization of CI/CD pipelines.