Score
Designs and implements validators, algorithms, and enforcement mechanisms that check structured patches or updates against full system state and global invariants, including declaring and enforcing authorized write regions and accepting or rejecting patches that violate constraints. Builds verification outputs (e.g., attestations or proofs) for valid updates and supports optics-style/structured-update representations to enable local validation, composition, and update propagation.
This work addresses the challenge of maintaining global consistency in LLM workflows that share structured state, where local updates can inadvertently violate system-wide invariants. The authors propose PatchOptic, a novel framework inspired by optics, which introduces a bidirectional access interface enabling safe local reads and writes through projected views, structured patches, and path-level footprint tracking. By integrating declarative region-based authorization, PatchOptic supports runtime validation, composable sub-workflows, and static certificate generation to ensure all updates adhere to global contracts. Evaluation on the PatchBench benchmark across 46 cases demonstrates that PatchOptic effectively blocks illicit updates—including contract violations and hidden-source attacks—while reducing information leakage and token overhead without compromising output quality.
This work addresses the frequent failures of electronic design automation (EDA) code generated by large language models (LLMs), which often arise from violations of implicit structural dependencies among design entities—such as invalid paths, missing preconditions, or API incompatibilities. To overcome the high latency and poor scalability of existing tool-in-the-loop debugging approaches, the authors propose a novel framework for reliable code generation that operates without runtime feedback. The key innovation lies in explicitly modeling structural dependencies as execution contracts and guiding a validator-driven synthesis process via a structural dependency graph. This approach integrates graph-conditioned retrieval, constraint generation, and staged pre-execution validation. Empirical results demonstrate a single-step task pass rate of 82.5%, an improvement in multi-step task success from 30.0% to 84.0%, over twofold reduction in tool invocations, and a validator precision of 93.3% (6.7% false positive rate).
Large language models (LLMs) lack verifiability and regulatory alignment when generating compliance-critical artifacts in safety-sensitive domains. Method: We propose Constraint-Guided Verifiable Generation (CVG), a framework featuring a Unified Meta-Model (UMM) for harmonizing heterogeneous regulatory texts; an Integrated Constraint Model (ICM) enabling dual-layer validation—structural (via GBNF/DFA) and semantic (via SHACL/SMT); and a synergistic prefix-safe decoding mechanism coupled with runtime automata and post-generation validators to embed auditable, traceable regulatory evidence chains. Contribution/Results: CVG innovatively integrates machine-verifiable certificates and violation-driven audit-and-repair directly into the generation pipeline. Evaluated on AUTOSAR automotive software and cross-border judicial workflows, CVG achieves 100% structural conformance, reduces manual correction effort by 72%, and seamlessly interoperates with existing Model-Driven Engineering (MDE) toolchains—delivering, for the first time, high-assurance, auditable, end-to-end compliant LLM-generated artifacts.
This work addresses a critical yet overlooked reliability issue in code generated by large language models (LLMs): despite passing compilation and unit tests, such code often fails in deployment due to structural inconsistencies—such as missing configurations, invalid imports, or omitted security controls—that evade detection by conventional CI/SAST tools. The paper introduces the “patchwork problem” to characterize these cross-module global defects, proposes an eight-category taxonomy specific to LLM-generated code, and formalizes structural consistency via invariants derived from a multidimensional code graph encompassing imports, calls, dependencies, configurations, and routing. Building on this foundation, the authors design a hybrid verification framework that integrates traditional static analysis with custom graph-based invariant checkers to precisely identify structural flaws invisible to existing tools. Empirical evaluation reveals that such defects are pervasive across major LLMs under diverse prompting strategies and exhibit distinct model-specific patterns.
Existing code-level formal verification tools scale poorly to large-scale software, while mainstream unit-level verification relies heavily on manual effort, often missing critical defects. This paper proposes the “Unit Proof Framework” research agenda—the first systematic definition of a unit verification paradigm supporting automated decoupling and independent verification of code units. Methodologically, it integrates formal verification, program analysis, modular verification, and automated toolchain design, with deep alignment to industrial development practices (e.g., AWS workflows). Its core contributions include: (1) establishing a scalable, engineering-friendly unit verification methodology; (2) characterizing a taxonomy of key technical challenges; (3) overcoming bottlenecks inherent in manual verification; and (4) significantly improving early detection of code-level defects. Collectively, this work lays the theoretical foundation and provides a practical technical pathway for building high-assurance, deployable automated verification infrastructure.
本文研究了四个去中心化构建包生态系统中的软件制品验证问题,通过定义独立验证模型和实现制品验证管道来解决因元数据缺失、隐式发布转换等问题导致的验证困难。
This study addresses the difficulty coding agents face in verifying whether programs satisfy specifications and their inability to effectively leverage expert diagnostic experience from verification failures. Building upon the executable semantics of the K framework, this work proposes a method that transforms expert diagnostics into reusable guidance, establishing a comprehensive verification pipeline encompassing specification generation, proof repair, and adequacy auditing. By designing paired clean and defective program packages, the effectiveness of the auditing mechanism is systematically evaluated. The proposed approach achieves a 164/164 pass rate on the HumanEval benchmark, with all defects accurately identified. Furthermore, experiments on KleverBench reveal optimization opportunities for guidance selection strategies under resource constraints. This research offers a novel paradigm for enhancing the formal verification capabilities of coding agents.
本文提出SecTB-RTL框架,通过31项任务和124个硬件安全回归测试,审计AI生成的RTL验证计划的有效性,发现仅满足提供者模式并不保证执行有效性。
This work addresses the problem of global inconsistency in multi-component intelligent agent releases, where local validation passes but cross-component relational integrity fails due to the absence of holistic consistency guarantees. To tackle this, we propose the Schema-SIP Relational Consistency (SIP-RC) framework—the first systematic approach to formally define and mitigate relational inconsistency faults in multi-component deployments. SIP-RC models release packages as graph structures and integrates schema documentation with product contract principles to enable cross-component relational verification. Key mechanisms include declarative–evidential linkage, decision authority scoping, provenance tracking of derived components, and byte-level consistency checks. Preliminary experiments demonstrate the feasibility of the proposed framework, offering a practical and actionable paradigm for ensuring relational consistency in intelligent agent releases.
论文提出一种框架明确指定基于大语言模型的程序修复实验设置,通过分析Defects4J和SWE-bench上的系统,解决了实验设置不透明导致的结果可比性问题。