multi-stage validation

Designs and implements multi-stage validation pipelines that cascade and progress from fast, deterministic programmatic checks (e.g., compilability and syntactic verification) to deeper equivalence and semantic evaluations (including model-/LLM-based tests) and to simulation or hardware-in-the-loop checks. Builds the orchestration and analysis that aggregates test outcomes, flags and routes failing cases for rejection, repair, or re-testing, and tunes thresholds and handoffs so the combined workflow catches both syntactic and semantic errors.

multi-stagevalidation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.26
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$207K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the frequent failures of electronic design automation (EDA) code generated by large language models (LLMs), which often arise from violations of implicit structural dependencies among design entities—such as invalid paths, missing preconditions, or API incompatibilities. To overcome the high latency and poor scalability of existing tool-in-the-loop debugging approaches, the authors propose a novel framework for reliable code generation that operates without runtime feedback. The key innovation lies in explicitly modeling structural dependencies as execution contracts and guiding a validator-driven synthesis process via a structural dependency graph. This approach integrates graph-conditioned retrieval, constraint generation, and staged pre-execution validation. Empirical results demonstrate a single-step task pass rate of 82.5%, an improvement in multi-step task success from 30.0% to 84.0%, over twofold reduction in tool invocations, and a validator precision of 93.3% (6.7% false positive rate).

EDA code generationreliable executionstructural dependencies

Automating a Complete Software Test Process Using LLMs: An Automotive Case Study

Feb 06, 2025
SW
Shuai Wang
🏛️ Chalmers University of Technology | Volvo Group

To address low testing efficiency in automotive API validation—caused by specification inconsistencies, protocol complexity, and high manual effort—this paper proposes the first end-to-end automated framework that deeply integrates large language models (LLMs) into the industrial-grade, full-lifecycle testing pipeline for automotive APIs. The framework employs task decomposition and multi-stage orchestration to enable requirement parsing, natural-language-driven test case generation, protocol-aware vehicle simulation interaction, and semantic result verification in a closed loop. Evaluated on over 100 real-world automotive APIs, it achieves a 92.3% test case generation accuracy and reduces human intervention by 87%. Notably, it establishes the first L3+-level fully autonomous, human-in-the-loop-free closed-loop testing capability—marking a significant departure from conventional script-based approaches that heavily rely on domain expertise.

Addresses complexity in API system alignment.Automates vehicle API testing using LLMs.Ensures stable, controlled testing workflow automation.

Natural language requirements are ill-suited for direct use in formal verification. Method: This work proposes an automated property generation framework integrating large language models (LLMs) with formal verification tools. It introduces an assertion generation mechanism extending beyond Linear Temporal Logic (LTL) to support numerical constraints and compositional system behavior modeling; integrates Claude 3.5 Sonnet with the ESBMC bounded model checker; and employs human-in-the-loop supervision to calibrate output quality. Contribution/Results: We first identify and characterize systematic impacts of LLM-induced model connection errors and numerical approximations on verification outcomes—reducing false positives and uncovering previously overlooked falsifiable scenarios. Evaluated on nine cyber-physical systems from Lockheed Martin, our approach achieves 46.5% verification accuracy—on par with NASA’s CoCoSim—while substantially lowering the barrier to formal verification and enhancing defect detection capability.

Automate deriving properties from natural language requirementsEnhance formal verification with LLM integrationReduce false positives in software verification

Model-Based Testing of an Intermediate Verifier Using Executable Operational Semantics

Aug 25, 2025
LL
Lidia Losavio
🏛️ USI Università della Svizzera italiana

Detecting subtle, specification-omitted bugs in Boogie—a widely used intermediate verification language—is challenging due to the incompleteness of existing formal models. Method: We propose BCC, a lightweight model-based testing technique grounded in executable operational semantics. BCC integrates the PLT Redex framework with a small, deterministic subset of Boogie’s operational semantics to automatically generate random programs; it then identifies bugs by comparing semantic simulation results against actual Boogie verification outcomes. Contribution/Results: BCC breaks from conventional reliance on full formal models by leveraging executable semantics to drive randomized testing—thereby efficiently exercising complex, non-canonical implementation paths in the toolchain. In evaluation, BCC generated 3 million test programs and uncovered completeness violations in 2% of them. These findings demonstrate BCC’s effectiveness and practicality for ensuring the reliability of verification tools themselves.

Detecting inconsistencies between operational semantics and verifier executionIdentifying potential bugs in formally verified verification toolsTesting Boogie intermediate verifier using model-based random generation

Workflow for Safe-AI

Mar 18, 2025
SV
Suzana Veljanovska
🏛️ ZHAW | Institute of Embedded Systems

Current AI model development for functional safety–critical domains (e.g., automotive and industrial control) lacks a systematic workflow that simultaneously ensures stability, certifiability, and adaptability. Method: This paper proposes a tool-certifiability–driven lightweight AI workflow paradigm. It introduces an extended ONNX-based, cross-stage verifiable AI model representation to unify modeling, verification, and deployment. The workflow integrates tool qualification (per ISO 26262/IEC 61508), the V-model development lifecycle, and static/dynamic AI verification techniques to enable end-to-end qualifiability of AI models in mixed-criticality systems. Contribution/Results: The approach significantly reduces tool qualification effort and supports reliable deployment across heterogeneous runtimes—including AUTOSAR and ROS 2. In representative use cases, it achieves 100% verification pass rate for model behavioral consistency.

Balance stability and adaptability in AI workflows.Develop a workflow for safe and dependable AI models.Ensure AI model validation from generation to deployment.

Latest Papers

What's happening recently
View more

This study addresses the limitations of existing SysML verification approaches, which are often tool-dependent and restricted to performance properties, lacking support for automated validation of behavioral and interface requirements. To overcome these shortcomings, this work proposes a tool-agnostic, automated verification workflow driven by SysML test cases, integrating UML Testing Profile and behavioral diagram constructs to enable unified validation of multidimensional attributes—including behavior, timing, and state responses. The methodology was developed through a mixed-methods research strategy combining literature review and stakeholder interviews, and its efficacy was empirically validated across two independent SysML toolchains. The approach not only transcends the constraints of conventional parametric methods but also enables automatic traceability of verification results back to the original model elements.

behavioral propertiesinterface propertiesmodel verification

This work addresses the limited generalization capability of large language models (LLMs) across hardware description languages, particularly due to the absence of a systematic evaluation framework for VHDL. We propose the first unified framework for LLM-based VHDL generation and evaluation, introducing an automated, verifiable Verilog-to-VHDL benchmark conversion pipeline. The resulting VHDLBench dataset comprises over 200 VHDL modules, each accompanied by complete testbenches. Integrating automated data synthesis, the VUnit/GHDL verification toolchain, and multi-model comparative analysis, our framework enables the first comprehensive assessment of LLM-generated VHDL code in terms of compilability, executability, and functional correctness. This study reveals critical challenges posed by VHDL-specific semantics and structural constructs, laying the groundwork for multilingual hardware design automation.

Hardware Description LanguagesLarge Language Modelsmodel generalization

This study addresses the frequent failure of large language models (LLMs) in software engineering tasks due to structurally invalid outputs—such as syntactic or formatting errors—that prevent correct parsing by downstream toolchains, even when the semantic content is accurate. The authors systematically evaluate the reliability of structured output generation across four representative tasks, categorizing errors into syntactic, structural, and semantic types. They propose a template-driven token-matching generation (TTMG) method that enforces structural consistency during autoregressive decoding. Experimental results demonstrate that TTMG nearly eliminates syntactic errors; however, structural and semantic errors remain prevalent. This work reveals, for the first time, that the fundamental bottleneck in structured output generation lies not merely in syntax but in the insufficient coordination between structural and semantic correctness, indicating that existing structural control mechanisms, while necessary, are insufficient without joint guarantees of both dimensions.

format compliancelarge language modelssoftware engineering

This work addresses the limitations of traditional simulation-based approaches in module-level fault analysis, which are often overly conservative and unable to accurately assess functional safety impacts. The authors propose SafeGen, a novel framework that integrates large language models (LLMs) with document-level hyperknowledge graphs (HyperKGs) to automatically extract verifiable specifications from design and safety documentation, generating semantically precise, design-aware functional safety assertions. By mapping gate-level faults to RTL and leveraging formal property verification (FPV), SafeGen enables semantic-level criticality classification for stuck-at and bridging faults, while supporting end-to-end traceable reasoning across specifications, assertions, and faults. Experimental evaluation on a field-oriented control (FOC) platform demonstrates that the generated assertions outperform those from existing LLM-based methods in quality and provide more semantically interpretable criticality assessments.

assertion generationautomotive chip designfault criticality

This study addresses a critical yet previously underexplored issue in large language model (LLM)-driven software development: the contamination of automatically generated tests by erroneous code. The authors systematically uncover and empirically validate this error propagation phenomenon, demonstrating that when tests are generated based on incorrect code within multi-step agent workflows—across diverse programming tasks and various prompting strategies, including chain-of-thought—the resulting tests exhibit significantly lower defect detection rates (14%) compared to independently generated tests (25%). These findings challenge the prevailing assumption that LLM-generated tests can serve as reliable, independent oracles, thereby highlighting the substantial risk of test bias in LLM-augmented development pipelines.

automated testing reliabilitycode-test consistencyerror propagation

Hot Scholars

SG

Shanghua Gao

Harvard University
Agentic AIAI for ScienceRepresentation learningGenerative modeling
HC

Henry C Woodruff

Head of The D-Lab, Department of Precision Medicine, Maastricht University
Medical ImagingMachine LearningBig DataRadiomics
SB

Simon Bing

TU Berlin
representation learningcausalityclimate
VB

Vahid Balazadeh

PhD in Computer Science - University of Toronto
Machine LearningCausality
LP

Lennart Purucker

PhD Student
Tabular DataBenchmarkingAutoMLMachine Learning