Score
Designs and implements agent-driven, multi-step verification systems that iteratively check and enforce category coverage, topology, geometric compatibility, and other correctness constraints when composing assemblies from candidate parts. Builds agentic verifiers and verification loops (including vision–language or perception–language loops) that filter and re-rank candidates, coordinate iterative checks across steps, and analyze assembled outputs for consistency and validity.
This work addresses the lack of structural correctness verification during the design phase in existing AI agent workflow platforms, which typically rely on runtime safeguards. The authors propose a workflow modeling approach centered on reusable building blocks and introduce, for the first time, a set of twelve structural rules. By leveraging graph-based representations and a rule engine, the method enables static, formal checks for compatibility and logical consistency at design time. Experimental evaluation demonstrates that the prototype system efficiently detects design violations on a dataset comprising 48 defective workflows and 168 structural variants, maintaining high detection accuracy even when tasks are split across multiple agents. This significantly enhances the reliability and maintainability of workflow designs.
Existing verification methods for AI agent systems struggle to assess the reliability of multi-step decision trajectories in dynamic environments. Through a systematic literature review of 257 studies, this work constructs a five-dimensional verification taxonomy encompassing behavioral, safety, temporal, regulatory, and multi-agent aspects. Analysis of case studies from healthcare, industrial automation, and intelligent transportation reveals critical gaps in current research, particularly concerning temporal validity, runtime evidence maintenance, regulatory interpretability, and assurance in open multi-agent settings. The study proposes a lifecycle-oriented verification agenda and outlines four key directions: bounded autonomy specifications, adversarial trajectory generation, runtime monitoring, and auditable evidence structures—collectively offering a pathway toward context-aware, trajectory-level trustworthy verification.
Existing agent-assisted design systems struggle to generate complex three-dimensional assemblies with moving parts, such as pistons or scissors. This work proposes AADvark, the first large language model–driven system capable of generating functional, movable CAD assemblies. By establishing a closed-loop pipeline that integrates code generation, model compilation, and visual feedback—augmented with a dedicated assembly constraint solver and a custom visual feedback mechanism—the system enables agents to directly reason about and synthesize dynamic structures with one or more degrees of freedom. This approach transcends the limitations of conventional static CAD modeling, successfully producing a variety of intricate movable assemblies while ensuring geometric and kinematic correctness through strong validation signals.
Formal verification ensures software correctness but suffers from high manual proof-writing costs, limiting its practical adoption. This work proposes a novel approach that integrates large language models with agent-guided tree search to enhance verification efficiency. By leveraging an iterative agent loop and mathlib retrieval, the method improves proof generation, further refined through a context-guided tree search architecture. Experimental results on 423 Lean specifications demonstrate a 95.0% verification success rate. Notably, the context-guided tree search significantly outperforms baselines on medium-difficulty problems with lower token consumption, while traditional agent iteration remains advantageous on the most challenging tasks, revealing distinct applicability regimes for different search strategies.
为解决LLM生成复杂程序的安全审查问题,MAGS通过多代理框架和形式化验证方法自动生成带安全保证的可执行程序。
This work addresses the vulnerability of existing large-scale multi-agent systems to failure in complex tasks due to error propagation and insufficient verification mechanisms. The authors propose a two-stage framework that automatically constructs and executes task-specific multi-agent systems from natural language instructions, incorporating dual verification mechanisms—during both construction and runtime. The approach decomposes tasks into directed acyclic graphs, defines input/output contracts, grounds knowledge via web search, and auto-generates prompts and tools. It further introduces a three-level error attribution scheme and intermediate output validation gating to enable targeted recovery strategies. Experimental results demonstrate that the method significantly outperforms strong baselines across programming, in-context learning, and open-ended reasoning tasks, consistently improving task success rates, error recovery capability, and workflow stability.
This study addresses the inability of LLM agents to autonomously report errors due to the lack of external verification, which typically forces reliance on budget exhaustion for termination. Drawing inspiration from the V-model in software engineering, this work proposes a hierarchical verification architecture. Its core innovation lies in decoupling zero-cost deterministic gating from optional LLM-based judging, enforced by a controller to achieve precise fault localization, proactive halting, and graceful degradation. Experiments on multi-hop question answering using the MuSiQue dataset demonstrate that, with an 8B-parameter model, the proposed architecture successfully answers eight out of ten questions while correctly refusing unanswerable ones, whereas the baseline fails entirely. Notably, most corrections are accomplished through the zero-cost gates, significantly enhancing overall system reliability.
This study addresses the critical issue of infinite agent loops (IAL) in large language model (LLM) agents, which can arise from missing or ineffective termination conditions during iterative execution, leading to resource exhaustion and denial-of-service. The work presents IAL-Scan, the first systematic approach to detecting IAL vulnerabilities, leveraging a unified intermediate representation to abstract heterogeneous agent code, constructing agent loop dependency graphs, and applying path reachability analysis to identify high-risk feedback loops. Evaluated on 6,549 open-source projects, IAL-Scan identified 74 potential IAL instances; manual validation confirmed 68 true positives across 47 projects, achieving a precision of 91.9%. This demonstrates IAL-Scan’s effectiveness in enabling cross-framework, high-precision discovery of IAL vulnerabilities.
本文探讨了循环工程在软件项目中的应用,通过设计系统自动启动和停止AI代理,并基于36,710个软件仓库的研究验证了其实际运行情况。
论文针对软件工程中的需求差距和模型差距问题,提出了一种利用部署证据来持续修订需求、模型或评估器的方法,以缩小这些差距。