generate and validate assertions

Designs and builds executable assertions and human-readable assertion descriptions by synthesizing checks from available oracles, specifications, or observed program behavior; this includes producing assertion candidates and encoding them as predicates or test-time checks. Validates and hardens those assertions by normalizing and verifying them, and by using coverage- and mutation-based techniques (and other analyses) to assess and improve their correctness, completeness, and robustness.

generateandvalidateassertions

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.53
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing large code models struggle to generate executable intermediate formal specifications, limiting precise verification and repair of program behavioral errors. This work proposes SpecCoder, a novel framework that focuses on generating executable inline assertions at critical program locations, thereby transforming static annotations into verifiable evidence. SpecCoder employs verification-guided training, fine-tuning the Qwen2.5-Coder series models using correct programs, behavioral mutants, and multi-round specification refinement trajectories. Evaluated on the HumanExec benchmark, SpecCoder substantially improves the correctness (+55.8%), completeness (+358.1%), and assertion validity (+26.6%) of inline specifications, significantly enhancing program verification and repair capabilities.

code LLMsexecutable assertionsformal specifications

Automated Test Generation from Program Documentation Encoded in Code Comments

Apr 29, 2025
GD
Giovanni Denaro
🏛️ University of Milano-Bicocca

Coverage-driven testing often fails to capture semantic behavior and misses deep-seated defects. To address this, this paper proposes a behavior-oriented test generation method leveraging code comments (e.g., Javadoc). First, functional specifications are extracted from natural-language comments and formally modeled as executable test objectives. Second, search-based test generation is integrated with context-aware assertion synthesis to automatically generate named, behaviorally meaningful test cases. This work is the first to explicitly transform code comments into executable test goals, thereby overcoming the semantic blindness of conventional coverage metrics and enabling precise, behavior-level testing. Evaluated on a benchmark of 118 Java classes, our approach significantly improves behavioral coverage and successfully detects multiple previously unknown defects that evade state-of-the-art coverage-based tools.

Detects unknown failures via contextualized test casesGenerates tests from code comments automaticallyTargets untested behaviors in coverage-driven approaches

This work addresses the challenge of poor generalizability in dynamically inferred specifications, which often stems from insufficient test coverage and necessitates extensive manual filtering. To mitigate this issue, the study introduces a novel integration of large language model (LLM)-generated counterexample tests into the dynamic inference pipeline. By leveraging tools such as SpecFuzzer to automatically validate the inferred assertions, the approach significantly improves precision without compromising recall. Experimental results demonstrate that the method effectively eliminates up to 11.68% of invalid assertions, achieving a maximum precision gain of 7% in specification inference. This advancement enhances both the accuracy and automation level of dynamic specification inference, reducing reliance on human intervention while maintaining robustness in inferred program specifications.

contract assertionsdynamic specification inferencefalse positives

Understanding and Characterizing Mock Assertions in Unit Tests

Mar 25, 2025
HZ
Hengcheng Zhu
🏛️ The Hong Kong University of Science and Technology | The University of Auckland | McGill University | Southern University of Science and Technology

Existing automated test generation techniques largely ignore mock assertions, and there is a lack of empirical evidence characterizing their usage, verification objectives, and fault-detection effectiveness. Method: This paper presents the first empirical study on mock assertions in unit tests, analyzing 4,652 test cases from 11 prominent open-source Java projects via combined static analysis and manual annotation to systematically identify and categorize mock assertion patterns. Contribution/Results: We find that mock assertions are primarily used to verify external dependency invocations, critical control-flow path execution, and internal side effects. They complement traditional assertions by uniquely enabling detection of side effects, control logic violations, and implicit state changes—capabilities inaccessible to conventional assertions. This work fills a critical gap in the empirical understanding of mock assertions and provides foundational insights and practical evidence to guide mock-aware test generation techniques.

Exploring mock assertions' effectiveness in fault detection and test generationInvestigating adoption and characteristics of mock assertions in Java projectsUnderstanding mock assertions' role in validating unobservable program behaviors

This work proposes a novel paradigm that bridges the long-standing divide between testing and formal verification in traditional software validation, enabling them to synergistically enhance both efficiency and quality. Grounded in Design by Contract, the approach leverages the counterexample generation capability of SMT solvers to transform formal verification tools into an integrated engine for automated testing and repair. Within a unified framework, the method simultaneously achieves three key objectives: automatic generation of test cases for faulty programs, construction of regression test suites with full coverage for correct programs, and correctness-guaranteed program repair. This represents the first integration of verification, testing, and repair into a single cohesive methodology.

automatic program repairformal proofsoftware maintenance

Latest Papers

What's happening recently
View more

This work addresses the challenges of ambiguity and validity verification in automatically translating natural language assertions into formal, executable specifications—a task traditionally reliant on error-prone manual effort. The paper proposes Monty, a novel framework that leverages large language models to generate candidate formal assertions and introduces an innovative combination of code-based testing and consistency scoring to automatically select high-quality translations. Evaluated on 541 tasks, Monty achieves up to a 20-percentage-point improvement in average precision over baseline approaches that directly use large language models for translation, substantially enhancing the accuracy and reliability of automated formalization.

autoformalizationformal contractsLLMs

This work addresses the high cost of manually writing formal specifications and the limitations of existing large language model (LLM)-based approaches that require white-box access to source code, thereby posing intellectual property and deployment constraints. The authors propose a black-box-driven method that leverages only test code and dynamic execution traces to generate candidate Java Modeling Language (JML) specifications via an LLM. These candidates are locally validated using bounded model checking, and an iterative feedback loop refines them based on verification outcomes. This approach is the first to enable fully automated formal specification generation without any access to the program’s internal structure. Evaluated on the SpecGenBench benchmark, it demonstrates that test-derived information effectively guides specification synthesis, while also highlighting critical challenges in checker compatibility and diagnostic feedback, substantially enhancing industrial applicability.

dynamic execution tracesformal specificationsLLM

Automatically generating high-quality, correct, and complete formal specifications—such as those in JML—remains a significant challenge: existing approaches often produce specifications that pass syntactic validation yet suffer from semantic inaccuracies or insufficient coverage. This work proposes VeriAct, a novel framework that introduces Spec-Harness, the first evaluation mechanism capable of precisely assessing both correctness and completeness of generated specifications. VeriAct further establishes the first verification-guided agent system, leveraging large language models within a closed-loop iterative process that integrates code execution, formal verification, and feedback signals to collaboratively synthesize and repair specifications. Experimental results demonstrate that VeriAct substantially outperforms current methods on two benchmarks, yielding specifications that not only satisfy verifiers but also achieve higher standards of semantic correctness and completeness.

completenesscorrectnessformal specification

This work addresses the limitation of large language models (LLMs) in generating SystemVerilog assertions that often fail to cover critical functional behaviors due to insufficient understanding of circuit designs. To overcome this, the authors propose CoverAssert, an iterative framework that, for the first time, integrates a syntax–semantics joint representation—based on abstract syntax trees and semantic feature clustering—with a functional coverage feedback mechanism. This approach maps assertions to natural language specifications and guides the LLM to prioritize generating assertions for uncovered scenarios. Experimental results on four open-source designs demonstrate that integrating AssertLLM with Spec2Assertion yields average improvements of 9.57%, 9.64%, and 15.69% in branch, statement, and toggle coverage, respectively, substantially enhancing verification completeness.

functional coverageIC designLLM

Traditional code auditing tools struggle to identify security vulnerabilities implied in natural language specifications and often produce false positives with ambiguous root causes. This work proposes SPECA, a framework that parses natural language specifications to extract explicit, typed security properties and integrates formal modeling with structured proof-based reasoning for vulnerability detection. SPECA’s key innovations include specification-aware precise auditing, unified cross-repository property comparison, and a traceable false positive attribution mechanism. Experimental evaluation demonstrates that SPECA successfully reproduces all known vulnerabilities on the Sherlock and RepoAudit benchmarks while uncovering multiple previously unknown flaws; furthermore, its false positives are systematically attributable to three distinct, actionable root causes amenable to improvement.

code-level auditingfalse positivesnatural-language specifications

Hot Scholars

HL

Huawei Li

Institute of Computing Technology, Chinese Academy of Sciences
computer engineering
XP

Xin Peng

East China University of Science and Technology
Artificial IntelligenceMachine LearningComplex Process Modeling
HS

Hailong Sun

Professor of Computer Science, Beihang University
Software EngineeringArtificial IntelligenceSoftware Systems
YW

Yuanli Wang

Boston University
Distributed SystemsMLSysLarge Language ModelsAgentic AI