Score
Designs and implements multi-stage validation pipelines that cascade and progress from fast, deterministic programmatic checks (e.g., compilability and syntactic verification) to deeper equivalence and semantic evaluations (including model-/LLM-based tests) and to simulation or hardware-in-the-loop checks. Builds the orchestration and analysis that aggregates test outcomes, flags and routes failing cases for rejection, repair, or re-testing, and tunes thresholds and handoffs so the combined workflow catches both syntactic and semantic errors.
This work addresses the frequent failures of electronic design automation (EDA) code generated by large language models (LLMs), which often arise from violations of implicit structural dependencies among design entities—such as invalid paths, missing preconditions, or API incompatibilities. To overcome the high latency and poor scalability of existing tool-in-the-loop debugging approaches, the authors propose a novel framework for reliable code generation that operates without runtime feedback. The key innovation lies in explicitly modeling structural dependencies as execution contracts and guiding a validator-driven synthesis process via a structural dependency graph. This approach integrates graph-conditioned retrieval, constraint generation, and staged pre-execution validation. Empirical results demonstrate a single-step task pass rate of 82.5%, an improvement in multi-step task success from 30.0% to 84.0%, over twofold reduction in tool invocations, and a validator precision of 93.3% (6.7% false positive rate).
To address low testing efficiency in automotive API validation—caused by specification inconsistencies, protocol complexity, and high manual effort—this paper proposes the first end-to-end automated framework that deeply integrates large language models (LLMs) into the industrial-grade, full-lifecycle testing pipeline for automotive APIs. The framework employs task decomposition and multi-stage orchestration to enable requirement parsing, natural-language-driven test case generation, protocol-aware vehicle simulation interaction, and semantic result verification in a closed loop. Evaluated on over 100 real-world automotive APIs, it achieves a 92.3% test case generation accuracy and reduces human intervention by 87%. Notably, it establishes the first L3+-level fully autonomous, human-in-the-loop-free closed-loop testing capability—marking a significant departure from conventional script-based approaches that heavily rely on domain expertise.
Natural language requirements are ill-suited for direct use in formal verification. Method: This work proposes an automated property generation framework integrating large language models (LLMs) with formal verification tools. It introduces an assertion generation mechanism extending beyond Linear Temporal Logic (LTL) to support numerical constraints and compositional system behavior modeling; integrates Claude 3.5 Sonnet with the ESBMC bounded model checker; and employs human-in-the-loop supervision to calibrate output quality. Contribution/Results: We first identify and characterize systematic impacts of LLM-induced model connection errors and numerical approximations on verification outcomes—reducing false positives and uncovering previously overlooked falsifiable scenarios. Evaluated on nine cyber-physical systems from Lockheed Martin, our approach achieves 46.5% verification accuracy—on par with NASA’s CoCoSim—while substantially lowering the barrier to formal verification and enhancing defect detection capability.
Detecting subtle, specification-omitted bugs in Boogie—a widely used intermediate verification language—is challenging due to the incompleteness of existing formal models. Method: We propose BCC, a lightweight model-based testing technique grounded in executable operational semantics. BCC integrates the PLT Redex framework with a small, deterministic subset of Boogie’s operational semantics to automatically generate random programs; it then identifies bugs by comparing semantic simulation results against actual Boogie verification outcomes. Contribution/Results: BCC breaks from conventional reliance on full formal models by leveraging executable semantics to drive randomized testing—thereby efficiently exercising complex, non-canonical implementation paths in the toolchain. In evaluation, BCC generated 3 million test programs and uncovered completeness violations in 2% of them. These findings demonstrate BCC’s effectiveness and practicality for ensuring the reliability of verification tools themselves.
Current AI model development for functional safety–critical domains (e.g., automotive and industrial control) lacks a systematic workflow that simultaneously ensures stability, certifiability, and adaptability. Method: This paper proposes a tool-certifiability–driven lightweight AI workflow paradigm. It introduces an extended ONNX-based, cross-stage verifiable AI model representation to unify modeling, verification, and deployment. The workflow integrates tool qualification (per ISO 26262/IEC 61508), the V-model development lifecycle, and static/dynamic AI verification techniques to enable end-to-end qualifiability of AI models in mixed-criticality systems. Contribution/Results: The approach significantly reduces tool qualification effort and supports reliable deployment across heterogeneous runtimes—including AUTOSAR and ROS 2. In representative use cases, it achieves 100% verification pass rate for model behavioral consistency.
This study addresses the limitations of existing SysML verification approaches, which are often tool-dependent and restricted to performance properties, lacking support for automated validation of behavioral and interface requirements. To overcome these shortcomings, this work proposes a tool-agnostic, automated verification workflow driven by SysML test cases, integrating UML Testing Profile and behavioral diagram constructs to enable unified validation of multidimensional attributes—including behavior, timing, and state responses. The methodology was developed through a mixed-methods research strategy combining literature review and stakeholder interviews, and its efficacy was empirically validated across two independent SysML toolchains. The approach not only transcends the constraints of conventional parametric methods but also enables automatic traceability of verification results back to the original model elements.
This work addresses the limited generalization capability of large language models (LLMs) across hardware description languages, particularly due to the absence of a systematic evaluation framework for VHDL. We propose the first unified framework for LLM-based VHDL generation and evaluation, introducing an automated, verifiable Verilog-to-VHDL benchmark conversion pipeline. The resulting VHDLBench dataset comprises over 200 VHDL modules, each accompanied by complete testbenches. Integrating automated data synthesis, the VUnit/GHDL verification toolchain, and multi-model comparative analysis, our framework enables the first comprehensive assessment of LLM-generated VHDL code in terms of compilability, executability, and functional correctness. This study reveals critical challenges posed by VHDL-specific semantics and structural constructs, laying the groundwork for multilingual hardware design automation.
This study addresses the frequent failure of large language models (LLMs) in software engineering tasks due to structurally invalid outputs—such as syntactic or formatting errors—that prevent correct parsing by downstream toolchains, even when the semantic content is accurate. The authors systematically evaluate the reliability of structured output generation across four representative tasks, categorizing errors into syntactic, structural, and semantic types. They propose a template-driven token-matching generation (TTMG) method that enforces structural consistency during autoregressive decoding. Experimental results demonstrate that TTMG nearly eliminates syntactic errors; however, structural and semantic errors remain prevalent. This work reveals, for the first time, that the fundamental bottleneck in structured output generation lies not merely in syntax but in the insufficient coordination between structural and semantic correctness, indicating that existing structural control mechanisms, while necessary, are insufficient without joint guarantees of both dimensions.
This work addresses the limitations of traditional simulation-based approaches in module-level fault analysis, which are often overly conservative and unable to accurately assess functional safety impacts. The authors propose SafeGen, a novel framework that integrates large language models (LLMs) with document-level hyperknowledge graphs (HyperKGs) to automatically extract verifiable specifications from design and safety documentation, generating semantically precise, design-aware functional safety assertions. By mapping gate-level faults to RTL and leveraging formal property verification (FPV), SafeGen enables semantic-level criticality classification for stuck-at and bridging faults, while supporting end-to-end traceable reasoning across specifications, assertions, and faults. Experimental evaluation on a field-oriented control (FOC) platform demonstrates that the generated assertions outperform those from existing LLM-based methods in quality and provide more semantically interpretable criticality assessments.
This study addresses a critical yet previously underexplored issue in large language model (LLM)-driven software development: the contamination of automatically generated tests by erroneous code. The authors systematically uncover and empirically validate this error propagation phenomenon, demonstrating that when tests are generated based on incorrect code within multi-step agent workflows—across diverse programming tasks and various prompting strategies, including chain-of-thought—the resulting tests exhibit significantly lower defect detection rates (14%) compared to independently generated tests (25%). These findings challenge the prevailing assumption that LLM-generated tests can serve as reliable, independent oracles, thereby highlighting the substantial risk of test bias in LLM-augmented development pipelines.