Score
Designing and implementing automated checkpoints that evaluate artifacts against explicit, measurable quality criteria to block or permit progression through a pipeline or lifecycle. Work includes specifying gate rules and thresholds, building enforcement and automation, integrating gates into workflows, and analyzing gate effectiveness and failure modes.
研究通过提示护栏和人工检查点两种方法,有效减少了多阶段LLM招聘流程中的信息伪造问题。
本文提出Agile-V Assurance Spine,通过权威源配置文件、工件绑定、风险适当独立性和时效性等方法解决工程生命周期中对代理输出的正当行动问题。
This work addresses the lack of closed-loop control in traditional software development lifecycles, which often fails to simultaneously ensure security, auditability, and highly reliable automation. The authors propose a deterministic autonomous control framework that models the lifecycle as a seven-stage automated pipeline, integrating Jira-based task orchestration, structured context, resource constraints, and human-review gating mechanisms to establish a secure closed loop. Key innovations include a state-contract-based collision locking mechanism, a degradation protocol for fallback operation, and a traceable control architecture. Implemented with 12,661 lines of Python code and 6,907 lines of versioned prompt specifications—including 101 exception handlers and 12 centralized locks—the system achieved a 100% success rate (95% CI [97.6%, 100%]) across 152 initial runs, producing over 795 artifacts. All 51 issues identified through adversarial review were fully resolved, with 60% of security tickets autonomously completed.
This study addresses the deficiency in Business Process Management (BPM) education, where overlooking AI’s central design role leaves students ill-equipped for process-level AI decision-making. To bridge this gap, we propose the “AI Decision Checkpoint” framework, which treats AI as a first-class citizen and explicitly distinguishes task-level automation from process-level value creation. By integrating large language models, retrieval-augmented generation, intelligent agents, and process modeling and mining techniques, the framework establishes a six-module closed-loop pedagogical system. Furthermore, it embeds AI evaluation and decision-making stages into a customer onboarding case study. Preliminary validation demonstrates that this approach effectively enhances students’ clarity in AI-integrated decision-making and their comprehension of process-level value creation.
Existing automated tool-calling systems often suffer from insufficient generalization due to model-centric designs and heavy reliance on prompting, leading to recurrent failures such as unsafe side effects, invalid parameters, uncontrolled retries, and sensitive data leakage. This work proposes a model-agnostic, policy-first framework for tool orchestration that enforces permission control prior to invocation, enhancing safety through explicit constraints, risk-aware gating, recovery mechanisms, and auditable explanations. Key contributions include a policy-first paradigm for tool workflows, a lightweight domain-specific language (DSL) for policies, a runtime execution engine, and a reproducible safety benchmark based on trajectory replay. In 225 controlled experiments, the strictest policy configuration achieved a violation prevention rate of 0.681, reduced retry amplification to 1.378, and attained a sensitive information leakage recall of 0.875, effectively quantifying the trade-off between safety and utility.
This study addresses the limited effectiveness of validators as hard gates in deployment pipelines, which is constrained by a “pass-if-not-run” skip defect that imposes a theoretical ceiling on detection capability. To investigate this, thirteen validators are evaluated based on statistical separability, employing Newcombe intervals, Fisher’s exact test, and multiple comparison corrections to rigorously quantify the distinction between execution states and static labels. The analysis reveals that execution bias accounts for approximately 84% of the detection ceiling, with only two checks demonstrating significant efficacy. Accordingly, this work proposes a new specification requiring evaluation records to carry rejection evidence, demonstrates that existing checklists are insufficient to guarantee gating quality, and establishes an execution-grounded standard for validator assessment.
This study addresses the silent failure of quality gates in release pipelines caused by missing data, which leads to unexecuted checks being erroneously interpreted as passed. To resolve this, we propose a three-valued decision logic that introduces an “indeterminate” state to distinguish genuine passes from missing data, thereby converting unexecuted checks into explicit failures. By integrating static analysis with CI/CD pipeline monitoring, we demonstrate that a single query suffices to remediate most defects of this nature. Experimental evaluation successfully identified nine latent defect cases and effectively halted erroneous green builds that had persisted for two weeks. These results validate both the reliability and falsifiability of the proposed remediation mechanism.
研究探讨了AI辅助软件工程中,通过分层监督(包括预防性、可执行性和人工监督)来应对传统代码审查等控制机制面临的压力。
This study addresses the absence of admission and assurance mechanisms for generative AI in safety-critical workflows by proposing a unit-level formal framework. Methodologically, it defines selective AI participation mechanisms and assurance records to decouple formal fidelity from validator soundness. It establishes multiple verification modes—including deterministic, statistically calibrated, human-in-the-loop, and CAD generation—alongside a non-uniform scale taxonomy, while introducing evidence obligation management. Experimental evaluation on a 17-unit wing spar analysis demonstrates that AI-generated CAD models successfully pass all 23 verification checks, achieving a gated instantiation of the AI-assisted workflow. This work thereby provides a verifiable paradigm for integrating AI into safety-critical engineering scenarios.
研究提出证据携带终止(ECT)方法,解决工具使用型代理何时停止的问题,通过绑定答案声明与有效范围内的追踪证据及确定性重放来确保安全终止。