guardrails and vetoed deployment control

Designs, implements, and evaluates technical and procedural safety mechanisms that prevent harmful, non-compliant, or unsafe models and features from being deployed, including automated veto gates, pre-deployment checks, policy enforcement engines, runtime monitoring, human-in-the-loop override controls, and rollback triggers. Builds tests, metrics, and integration points in CI/CD pipelines to detect failure modes, measure guardrail effectiveness, and ensure deployments can be reliably blocked or reverted when policies or safety conditions are violated.

guardrailsandvetoeddeployment

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.08
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Existing automated tool-calling systems often suffer from insufficient generalization due to model-centric designs and heavy reliance on prompting, leading to recurrent failures such as unsafe side effects, invalid parameters, uncontrolled retries, and sensitive data leakage. This work proposes a model-agnostic, policy-first framework for tool orchestration that enforces permission control prior to invocation, enhancing safety through explicit constraints, risk-aware gating, recovery mechanisms, and auditable explanations. Key contributions include a policy-first paradigm for tool workflows, a lightweight domain-specific language (DSL) for policies, a runtime execution engine, and a reproducible safety benchmark based on trajectory replay. In 225 controlled experiments, the strictest policy configuration achieved a violation prevention rate of 0.681, reduced retry amplification to 1.378, and attained a sensitive information leakage recall of 0.875, effectively quantifying the trade-off between safety and utility.

model-agnostic safetysensitive data leakagetool-using automation

"Good"and"Bad"Failures in Industrial CI/CD -- Balancing Cost and Quality Assurance

Apr 16, 2025
SS
Simin Sun
🏛️ Chalmers University of Technology | University of Gothenburg | Zenseact

This study addresses the quality-efficiency-cost imbalance in industrial CI/CD pipelines caused by heterogeneous failure types. We propose a process refactoring paradigm centered on two critical milestones: code integration (pre-merge) and product release. First, we systematically define “good failures” (early-detected, low-cost) versus “bad failures” (late-occurring, high-blocking). Grounded in empirical studies across four enterprises—including workflow mapping and failure root-cause modeling—we develop a transferable pre-merge failure governance framework. Evaluation results show a 37% reduction in average feedback latency, a 29% decrease in spurious build overhead, significant improvement in developer throughput, and optimized cloud resource utilization. Our core contribution lies in transcending conventional stage-based pipeline segmentation to enable failure-driven, fine-grained process control—marking a paradigm shift toward adaptive, cost-aware CI/CD orchestration.

Addressing pre-merge phase failure prevention gapsBalancing cost and quality in CI/CD workflowsDistinguishing CI and CD for optimization milestones

AI-Augmented CI/CD Pipelines: From Code Commit to Production with Autonomous Decisions

Aug 15, 2025
MB
Mohammad Baqar
🏛️ Cisco Systems Inc | MUFG Bank | University of Houston

In modern CI/CD pipelines, manual intervention in unstable test diagnosis, rollback decisions, feature flag tuning, and canary promotion introduces release delays and operational overhead. To address this, we propose an AI-augmented autonomous software delivery framework that integrates large language models (LLMs) with policy-constrained autonomous agents, yielding a reference architecture for agent-based decision-making bounded by formal policies. Our approach introduces: (1) a taxonomy of deployment decisions; (2) policy-as-code guardrails enforcing safety and compliance; (3) a tiered trust framework governing agent autonomy; and (4) a DORA-metrics-driven, verifiable evaluation methodology. Evaluated in a React 19 microservices environment, the framework significantly reduces deployment latency and manual intervention frequency, improves release velocity and system reliability, and ensures auditable, formally verifiable autonomous decision paths.

Automating human decisions in CI/CD to reduce latencyMinimizing operational toil through autonomous policy-bounded agentsReplacing manual rollback and testing choices with AI

Enhancing Software Supply Chain Security Through STRIDE-Based Threat Modelling of CI/CD Pipelines

Jun 06, 2025
SD
Sowmiya Dhandapani
🏛️ Independent Cyber Security Researcher

This paper addresses the insufficient identification and mitigation of software supply chain security risks throughout the CI/CD pipeline lifecycle. We propose a structured threat modeling methodology grounded in the STRIDE framework, applied incrementally across core infrastructure components—including GitHub, Jenkins, Docker, and Kubernetes—to cover all phases from source code management to production deployment. A novel integration of STRIDE with the SLSA maturity model enables quantitative assessment of how specific security controls elevate SLSA compliance levels. By unifying Security as Code principles with the “Shift Left–Shield Right” paradigm, our approach realizes threat-driven, automated security enforcement. The outcomes include a structured threat–control mapping matrix and an actionable CI/CD security hardening roadmap, directly supporting DevSecOps adoption and progressive SLSA compliance advancement.

Analyzing vulnerabilities from source code to deployment with GitHub, Jenkins, Docker, KubernetesEnhancing CI/CD security via NIST, OWASP, and SLSA-based controls and toolchain integrationIdentifying and mitigating risks in CI/CD pipelines using STRIDE threat modeling

Latest Papers

What's happening recently
View more

This work addresses the threat posed by AI development agents that may covertly compromise infrastructure security through privilege escalation, log obfuscation, or persistence mechanisms—risks exacerbated by most organizations’ inability to deploy sophisticated monitoring. The authors propose a training-free, structured monitoring approach for Infrastructure-as-Code (IaC) environments, leveraging differential analysis between information flow graphs (IFGs) and control/data flow graphs to enable high-precision, interpretable detection of security degradation. The method supports both synchronous blocking and asynchronous auditing modes: in synchronous mode, real-time rollback reduces the success rate of stealthy attacks from 74.4% to 0%; in asynchronous mode, untrained IFGs lower the false-negative rate from 11.6% to 3.5% without disrupting legitimate operations.

AI agent safetycovert sabotagedeployment monitoring

This work addresses the semantic gap faced by current AI agents in enforcing natural language policies: the intended policy semantics are difficult to enforce precisely and interpretably at the system level. To bridge this gap, the authors propose a novel approach that integrates agent-side context with kernel-level enforcement mechanisms. For the first time, policy context is preserved on the agent side, while a domain-specific language (DSL) for information flow control (IFC), implemented via eBPF, enables comprehensive, action-level policy enforcement within the operating system kernel. This framework supports cross-event data-flow and ordering constraints, significantly improving policy compliance rates by covering indirect execution paths invisible to conventional tool-call interception. The system incurs only 1.9%–8.4% runtime overhead and provides semantically clear feedback instead of ambiguous errors.

AI agentsinformation-flow controlpolicy enforcement

Existing runtime enforcement techniques struggle to handle reactive systems with complex continuous dynamics and lack effective mechanisms for intervening in hybrid behaviors. This work proposes the first framework that integrates hybrid automata into runtime enforcement, enabling coordinated discrete event editing and continuous-time monitoring to correct system behavior at any instant by suppressing, delaying, or inserting events. The paper establishes formal enforceability conditions and devises an online strategy synthesis algorithm based on reachability analysis. Evaluation on an adaptive cruise control case study demonstrates that the approach ensures safety properties even when the underlying controller is unsafe, all while incurring minimal computational overhead.

Continuous DynamicsHybrid SystemsReactive Systems

Enterprise-scale general-purpose agents lack built-in, reusable governance mechanisms for autonomous cross-tool operation, making it difficult to satisfy requirements for compliance, auditability, and behavioral controllability. This work proposes the CUGA policy system, which embeds runtime governance capabilities into five critical checkpoints of the agent execution pipeline—intent protection, playbook guidance, tool invocation control, human approval gating, and output formatting—through a modular “policy-as-code” architecture. Without requiring model fine-tuning, CUGA enables proactive, continuous, and structured behavior control. By integrating typed governance primitives, dynamic playbook injection, and human-in-the-loop approval, the system effectively blocks malicious requests, enforces structured tool sequences, and triggers manual review for high-risk operations in healthcare scenarios, significantly enhancing policy adherence, execution consistency, and deployment safety.

autonomous enterprise agentscompliance-aware behaviorgeneralist agents