Score
Designs and implements verification systems that compare and reconcile outputs and behaviors across multiple models and heterogeneous validators, including LLM-based validators, to detect discrepancies in completeness, parameter accuracy, and execution order. Builds consensus and workflow-level validation pipelines plus structured self-correction loops that generate targeted feedback and re-run or repair actions to converge on consistent, correct results.
This study addresses the limitations of existing SysML verification approaches, which are often tool-dependent and restricted to performance properties, lacking support for automated validation of behavioral and interface requirements. To overcome these shortcomings, this work proposes a tool-agnostic, automated verification workflow driven by SysML test cases, integrating UML Testing Profile and behavioral diagram constructs to enable unified validation of multidimensional attributes—including behavior, timing, and state responses. The methodology was developed through a mixed-methods research strategy combining literature review and stakeholder interviews, and its efficacy was empirically validated across two independent SysML toolchains. The approach not only transcends the constraints of conventional parametric methods but also enables automatic traceability of verification results back to the original model elements.
This study addresses the challenge of specification scarcity in deductive verification of Java programs by proposing an LLM-driven, annotation-based JML specification generation and formal verification closed-loop framework. Methodologically, it treats LLMs (e.g., CodeLlama, GPT) as unreliable oracles; task-oriented prompt engineering elicits candidate specifications, which are then subjected to automated provability checking and iterative refinement via an SMT-solver–driven verifier (OpenJML/KeY). The key contribution is the first instantiation of an LLM–verifier collaborative closed-loop paradigm, enabling deep synergy between specification synthesis and formal verification. Evaluated on standard Java benchmark suites, 87% of LLM-generated specifications pass automatic verification, with 62% achieving full provable correctness—marking a substantial improvement in both reliability and practical applicability of generated specifications.
This work addresses the challenge of semantic drift and non-executable outputs in enterprise-grade large language model (LLM) assistants for code generation and business analytics, which typically rely on manual validation due to the absence of built-in verification mechanisms. To overcome this limitation, the authors propose two novel automated validation frameworks—Q* and Feedback+—that, for the first time, integrate reverse-translation semantic matching and code execution feedback loops into a conversational analytics system. By adopting a generator–discriminator architecture, the approach enables a paradigm shift from user-dependent validation to system-level self-verification. Experimental results on the Spider, Bird, and GSM8K benchmarks demonstrate significant reductions in error rates and task completion time, confirming the effectiveness and practical utility of the proposed methods.
This study addresses the lack of effective validation methods for semi-formal blueprints in early-stage software product line engineering, which often leads to undetected structural and constraint-related errors in feature models. For the first time, it systematically evaluates the capability of large language models (LLMs) in feature model analysis tasks by leveraging twelve state-of-the-art LLMs and sixteen standard analytical operations that integrate structural parsing with constraint reasoning. Performance is benchmarked against the solver-based tool FLAMA. Results demonstrate that reasoning-optimized models—such as Grok 4 Fast Reasoning and Gemini 2.5 Pro—achieve average accuracies of 88–89%, approaching the performance of formal solvers. These findings substantiate the feasibility and practical potential of LLMs as lightweight, early-stage validation tools for feature model verification.
Current benchmarks for mathematical reasoning predominantly rely on answer matching, which fails to assess the logical correctness of solution processes. This work proposes a hybrid verification pipeline that integrates automated and interactive validation by leveraging structured prompting to guide large language models in generating verifiable solutions. The framework supports both formal and informal reasoning and interfaces with proof assistants such as Lean 4, enabling even small-scale models (≤8B parameters) to participate effectively in collaborative verification. Through a multi-agent architecture and advanced prompt engineering, the approach substantially reduces false positive rates. Experimental results demonstrate high verification accuracy across multiple datasets, and the codebase along with deployment guidelines has been publicly released.
This work addresses the challenge of verifying the correctness of large language model (LLM) outputs and providing fine-grained feedback by introducing a general-purpose, training-free verification framework. The approach conceptualizes verification as a new dimension for enhancing agent capabilities, generating continuous scores through the expected value of token logits distributions. It supports refined scoring granularity, repeated evaluation, and criterion decomposition, enabling scalable, highly accurate, and well-calibrated verification. By integrating a cost-efficient ranking algorithm with a reinforcement learning–oriented dense feedback mechanism, the framework achieves state-of-the-art performance across multiple benchmarks: Terminal-Bench V2 (86.5%), SWE-Bench Verified (78.2%), RoboRewardBench (87.4%), and MedAgentBench (73.3%), substantially improving sample efficiency in code monitoring and reinforcement learning settings.
This work addresses the frequent failures of electronic design automation (EDA) code generated by large language models (LLMs), which often arise from violations of implicit structural dependencies among design entities—such as invalid paths, missing preconditions, or API incompatibilities. To overcome the high latency and poor scalability of existing tool-in-the-loop debugging approaches, the authors propose a novel framework for reliable code generation that operates without runtime feedback. The key innovation lies in explicitly modeling structural dependencies as execution contracts and guiding a validator-driven synthesis process via a structural dependency graph. This approach integrates graph-conditioned retrieval, constraint generation, and staged pre-execution validation. Empirical results demonstrate a single-step task pass rate of 82.5%, an improvement in multi-step task success from 30.0% to 84.0%, over twofold reduction in tool invocations, and a validator precision of 93.3% (6.7% false positive rate).
This work addresses the challenges of ensuring correctness and accurately constructing formal specifications when large language models generate code from natural language. We propose a verifiable code generation approach that integrates hierarchical prompting with verification feedback. To support this, we introduce the NL2VC-60 dataset, which leverages Dafny formal specifications and the uDebug platform to prevent vacuous verification. Furthermore, we design a self-repair prompting mechanism guided by structural signatures and verifier feedback. Experimental results demonstrate that our method substantially enhances both verifiability and functional correctness of code generated by open-source large models: Gemma-4-31B achieves a verification success rate of 90.91%, while GPT-OSS-120B improves from 0% to 81.82% under signature-guided prompting, marking the first systematic validation of open-source models’ potential to produce high-assurance code for complex algorithmic tasks.
This study addresses the common omission in current large language model evaluations of code generation—the iterative refinement process inherent in real-world programming and the models’ capacity for self-correction using feedback. The authors propose a novel framework that leverages execution-based feedback, such as compilation errors and test failures, to systematically investigate how reasoning and non-reasoning models utilize such signals across multiple programming languages. Through multidimensional categorization of code failures and extensive cross-model, cross-language experiments, they demonstrate that reasoning models consistently improve over iterations and significantly outperform non-reasoning counterparts. While syntactic and runtime errors prove relatively amenable to correction, logical and algorithmic errors remain challenging, thereby delineating the current limits of feedback-driven repair mechanisms.
This work addresses the vulnerability of large language model (LLM) agents to non-atomic failures—such as timeouts, delayed visibility, and partial state updates—in real-world systems when executing multi-step tasks via external tools, which often leads to redundant operations or task failure. To enhance robustness without modifying the underlying LLM, the authors propose a lightweight, verification-aware tool-wrapping mechanism that uniquely integrates post-condition validation with idempotency keys and introduces pre-retry verification logic. Experimental results demonstrate that this approach significantly reduces redundant tool invocations across diverse task templates while maintaining success rates comparable to baseline methods, thereby substantially improving the reliability of LLM agents operating in environments prone to non-atomic failures.