Score
Designs, implements, and evaluates technical and procedural safety mechanisms that prevent harmful, non-compliant, or unsafe models and features from being deployed, including automated veto gates, pre-deployment checks, policy enforcement engines, runtime monitoring, human-in-the-loop override controls, and rollback triggers. Builds tests, metrics, and integration points in CI/CD pipelines to detect failure modes, measure guardrail effectiveness, and ensure deployments can be reliably blocked or reverted when policies or safety conditions are violated.
To address runtime safety failures in large language model (LLM)-based agents—arising from autonomy and non-determinism in multi-step decision-making, goal planning, and tool invocation—this paper proposes a dynamic, hierarchical, cross-phase safety assurance framework. Methodologically, it innovatively adapts the Swiss Cheese Model to establish an AI safety reference architecture; introduces the first runtime protection taxonomy spanning quality attributes, pipeline stages, and architectural components; and formalizes an AI safety-by-design software architecture paradigm. Through systematic literature review, architectural modeling, and multi-layered defense design, the framework enables real-time monitoring and intervention over critical artifacts—including goals, plans, and tools. The contributions include a reusable classification system and a structured design guideline, collectively supporting the development of robust, verifiable, and evolvable safety-critical systems for foundation model agents. (149 words)
Existing automated tool-calling systems often suffer from insufficient generalization due to model-centric designs and heavy reliance on prompting, leading to recurrent failures such as unsafe side effects, invalid parameters, uncontrolled retries, and sensitive data leakage. This work proposes a model-agnostic, policy-first framework for tool orchestration that enforces permission control prior to invocation, enhancing safety through explicit constraints, risk-aware gating, recovery mechanisms, and auditable explanations. Key contributions include a policy-first paradigm for tool workflows, a lightweight domain-specific language (DSL) for policies, a runtime execution engine, and a reproducible safety benchmark based on trajectory replay. In 225 controlled experiments, the strictest policy configuration achieved a violation prevention rate of 0.681, reduced retry amplification to 1.378, and attained a sensitive information leakage recall of 0.875, effectively quantifying the trade-off between safety and utility.
This study addresses the quality-efficiency-cost imbalance in industrial CI/CD pipelines caused by heterogeneous failure types. We propose a process refactoring paradigm centered on two critical milestones: code integration (pre-merge) and product release. First, we systematically define “good failures” (early-detected, low-cost) versus “bad failures” (late-occurring, high-blocking). Grounded in empirical studies across four enterprises—including workflow mapping and failure root-cause modeling—we develop a transferable pre-merge failure governance framework. Evaluation results show a 37% reduction in average feedback latency, a 29% decrease in spurious build overhead, significant improvement in developer throughput, and optimized cloud resource utilization. Our core contribution lies in transcending conventional stage-based pipeline segmentation to enable failure-driven, fine-grained process control—marking a paradigm shift toward adaptive, cost-aware CI/CD orchestration.
研究探讨了AI辅助软件工程中,通过分层监督(包括预防性、可执行性和人工监督)来应对传统代码审查等控制机制面临的压力。
In modern CI/CD pipelines, manual intervention in unstable test diagnosis, rollback decisions, feature flag tuning, and canary promotion introduces release delays and operational overhead. To address this, we propose an AI-augmented autonomous software delivery framework that integrates large language models (LLMs) with policy-constrained autonomous agents, yielding a reference architecture for agent-based decision-making bounded by formal policies. Our approach introduces: (1) a taxonomy of deployment decisions; (2) policy-as-code guardrails enforcing safety and compliance; (3) a tiered trust framework governing agent autonomy; and (4) a DORA-metrics-driven, verifiable evaluation methodology. Evaluated in a React 19 microservices environment, the framework significantly reduces deployment latency and manual intervention frequency, improves release velocity and system reliability, and ensures auditable, formally verifiable autonomous decision paths.
This paper addresses the insufficient identification and mitigation of software supply chain security risks throughout the CI/CD pipeline lifecycle. We propose a structured threat modeling methodology grounded in the STRIDE framework, applied incrementally across core infrastructure components—including GitHub, Jenkins, Docker, and Kubernetes—to cover all phases from source code management to production deployment. A novel integration of STRIDE with the SLSA maturity model enables quantitative assessment of how specific security controls elevate SLSA compliance levels. By unifying Security as Code principles with the “Shift Left–Shield Right” paradigm, our approach realizes threat-driven, automated security enforcement. The outcomes include a structured threat–control mapping matrix and an actionable CI/CD security hardening roadmap, directly supporting DevSecOps adoption and progressive SLSA compliance advancement.
This work addresses the threat posed by AI development agents that may covertly compromise infrastructure security through privilege escalation, log obfuscation, or persistence mechanisms—risks exacerbated by most organizations’ inability to deploy sophisticated monitoring. The authors propose a training-free, structured monitoring approach for Infrastructure-as-Code (IaC) environments, leveraging differential analysis between information flow graphs (IFGs) and control/data flow graphs to enable high-precision, interpretable detection of security degradation. The method supports both synchronous blocking and asynchronous auditing modes: in synchronous mode, real-time rollback reduces the success rate of stealthy attacks from 74.4% to 0%; in asynchronous mode, untrained IFGs lower the false-negative rate from 11.6% to 3.5% without disrupting legitimate operations.
This work addresses the semantic gap faced by current AI agents in enforcing natural language policies: the intended policy semantics are difficult to enforce precisely and interpretably at the system level. To bridge this gap, the authors propose a novel approach that integrates agent-side context with kernel-level enforcement mechanisms. For the first time, policy context is preserved on the agent side, while a domain-specific language (DSL) for information flow control (IFC), implemented via eBPF, enables comprehensive, action-level policy enforcement within the operating system kernel. This framework supports cross-event data-flow and ordering constraints, significantly improving policy compliance rates by covering indirect execution paths invisible to conventional tool-call interception. The system incurs only 1.9%–8.4% runtime overhead and provides semantically clear feedback instead of ambiguous errors.
Existing runtime enforcement techniques struggle to handle reactive systems with complex continuous dynamics and lack effective mechanisms for intervening in hybrid behaviors. This work proposes the first framework that integrates hybrid automata into runtime enforcement, enabling coordinated discrete event editing and continuous-time monitoring to correct system behavior at any instant by suppressing, delaying, or inserting events. The paper establishes formal enforceability conditions and devises an online strategy synthesis algorithm based on reachability analysis. Evaluation on an adaptive cruise control case study demonstrates that the approach ensures safety properties even when the underlying controller is unsafe, all while incurring minimal computational overhead.
Enterprise-scale general-purpose agents lack built-in, reusable governance mechanisms for autonomous cross-tool operation, making it difficult to satisfy requirements for compliance, auditability, and behavioral controllability. This work proposes the CUGA policy system, which embeds runtime governance capabilities into five critical checkpoints of the agent execution pipeline—intent protection, playbook guidance, tool invocation control, human approval gating, and output formatting—through a modular “policy-as-code” architecture. Without requiring model fine-tuning, CUGA enables proactive, continuous, and structured behavior control. By integrating typed governance primitives, dynamic playbook injection, and human-in-the-loop approval, the system effectively blocks malicious requests, enforces structured tool sequences, and triggers manual review for high-risk operations in healthcare scenarios, significantly enhancing policy adherence, execution consistency, and deployment safety.
研究提出GuardrailLoop测试平台,通过固定策略、计算限制和崩溃恢复方法解决自改进代理流程中的审计问题。