pattern validation and testing

Designs and implements validators and test suites that verify that a pattern—such as a structural, behavioral, or syntactic template—meets its specification and behaves correctly across representative and adversarial inputs by creating criteria, test cases, and automated checks for correctness, robustness, and edge cases. Analyzes test results, failure modes, coverage gaps, and performance implications to report deviations and recommend refinements or fixes to the pattern or its specification.

patternvalidationandtesting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.27
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing testing approaches for regular expression engines are hindered by syntactic discrepancies across dialects and inefficient fuzzing strategies, limiting their ability to uncover defects effectively. This work proposes ReTest, a novel framework that integrates grammar-aware fuzzing with 16 Kleene algebra–based homomorphic metamorphic relations, enabling the first systematic, cross-dialect, dependency-free testing methodology. By modeling regex semantics and employing coverage-guided exploration, ReTest achieves threefold higher edge coverage than state-of-the-art techniques on PCRE and discovers three previously unknown memory-safety vulnerabilities, substantially advancing the reliability validation of regex engines.

differential testingfuzzingmetamorphic testing

Covering All the Bases: Type-Based Verification of Test Input Generators

Apr 06, 2023
ZZ
Zhe-Wei Zhou
🏛️ Purdue University | Indian Institute of Technology Hyderabad

Verifying coverage completeness of input generators in property-based testing remains challenging. Method: This paper proposes a static verification approach based on a “must-style” refinement type system, reformulating conventional “may-produce” type semantics into “must-produce” semantics. It formally defines full coverage for higher-order functions and inductive data types, enabling fully automated verification of generator completeness. Contribution/Results: To our knowledge, this is the first refinement type system provably guaranteeing generation of all inputs satisfying both type and constraint specifications. Experimental evaluation demonstrates substantial improvements in detecting coverage gaps across diverse complex generators, while significantly reducing manual verification effort.

Ensures generators produce all required input valuesSupports polymorphism for real-world PBT frameworksValidates test generator coverage using refinement types

This work addresses the limited adoption of formal verification, which often requires expert-written annotations such as preconditions, postconditions, and loop invariants. To overcome this barrier, the authors propose a novel approach that leverages large language models (LLMs) in conjunction with assertions from test cases as static oracles to automatically generate Dafny verification annotations from code annotated with natural language comments. The method features an iterative refinement process guided by verifier feedback over multiple rounds and uniquely integrates multi-model LLM collaboration with a closed-loop verifier feedback mechanism. A VS Code plugin was developed to support practical deployment. Evaluated on 110 Dafny programs, the approach achieves a 98.2% annotation correctness rate within at most eight repair iterations. Empirical results highlight that proof-assistant-style annotation remains a key challenge for LLMs, while user feedback on the plugin was notably positive.

Dafnyformal specificationLLMs

A Design Recipe and Recipe-Based Errors for Regular Expressions

Aug 05, 2025
MT
Marco T. Morazán
🏛️ Seton Hall University | Axoni | Penguin Random House | University of Washington

Students commonly struggle with constructing and debugging regular expressions due to cognitive overload and lack of systematic guidance. Method: This paper proposes a pedagogically oriented, structured design support framework that integrates the “design recipe” methodology to decompose regex construction into incremental, actionable steps; introduces a novel, stage-based error classification model that generates concise, jargon-free, non-prescriptive feedback; and incorporates a lightweight unit-testing shorthand syntax to lower verification barriers. Contribution/Results: Classroom deployment demonstrates significant improvements in students’ correctness rates and conceptual understanding during regex construction. Two authentic debugging case studies confirm the framework’s usability and pedagogical effectiveness, validating its capacity to enhance both learning efficiency and diagnostic accuracy in regex education.

Creating customized error messaging systemDeveloping shorthand syntax for unit testsProviding design support for regular expressions

Re-evaluation of Logical Specification in Behavioural Verification

May 23, 2025
JS
Jakub Semczyszyn
🏛️ AGH University of Krakow

Theorem provers exhibit unstable performance, poor scalability, and low reproducibility in behavioral verification—particularly for real-time and safety-critical software. Method: We systematically reproduce and extend existing benchmarks to construct an empirical evaluation framework featuring diverse behavioral models and logical specifications, enabling rigorous assessment of robustness, scalability, and reproducibility across mainstream theorem provers. Contribution/Results: Our study is the first to empirically establish strong correlations between irregular solver performance and structural problem characteristics—specifically temporal constraint density and state-space distribution. Leveraging these insights, we propose adaptive heuristic strategies and a self-optimizing solver architecture. The approach delivers measurable stability improvements for just-in-time verification in CI/CD pipelines and AI-augmented IDEs, significantly enhancing the practicality and trustworthiness of automated logical verification in high-assurance software development.

Enhancing stability of automated reasoning for safety-critical software verificationIdentifying performance irregularities in theorem provers across problem structuresValidating robustness and scalability of automated logical specification methods

Latest Papers

What's happening recently
View more

This work addresses the high cost of manually writing formal specifications and the limitations of existing large language model (LLM)-based approaches that require white-box access to source code, thereby posing intellectual property and deployment constraints. The authors propose a black-box-driven method that leverages only test code and dynamic execution traces to generate candidate Java Modeling Language (JML) specifications via an LLM. These candidates are locally validated using bounded model checking, and an iterative feedback loop refines them based on verification outcomes. This approach is the first to enable fully automated formal specification generation without any access to the program’s internal structure. Evaluated on the SpecGenBench benchmark, it demonstrates that test-derived information effectively guides specification synthesis, while also highlighting critical challenges in checker compatibility and diagnostic feedback, substantially enhancing industrial applicability.

dynamic execution tracesformal specificationsLLM

This study addresses the difficulty coding agents face in verifying whether programs satisfy specifications and their inability to effectively leverage expert diagnostic experience from verification failures. Building upon the executable semantics of the K framework, this work proposes a method that transforms expert diagnostics into reusable guidance, establishing a comprehensive verification pipeline encompassing specification generation, proof repair, and adequacy auditing. By designing paired clean and defective program packages, the effectiveness of the auditing mechanism is systematically evaluated. The proposed approach achieves a 164/164 pass rate on the HumanEval benchmark, with all defects accurately identified. Furthermore, experiments on KleverBench reveal optimization opportunities for guidance selection strategies under resource constraints. This research offers a novel paradigm for enhancing the formal verification capabilities of coding agents.

coding agentsformal specificationprogram correctness

This study addresses the lack of deterministic verification of user intent in formal specification generation, which risks producing proofs grounded in erroneous specifications. To this end, it constructs a unified dataset and multidimensional evaluation framework encompassing formal validity, similarity, and behavioral adequacy. The work proposes a strategy distinguishing input acceptance from output constraints to clarify the evidential scope of metrics, and integrates an LLM agent workflow, the Lean theorem prover, and Generalized Tree Edit Distance (GTED) to enable automated evaluation. The findings reveal the limitations of single similarity metrics and the impact of metric coverage on ranking outcomes, demonstrating that perfect postcondition scores may obscure deficiencies in input contracts. Ultimately, this research provides a systematic benchmark for evaluating formal specifications.

Formal Specification GenerationIntent AlignmentSpecGen Evaluation

This work addresses the limitation of conventional execution coverage in UI component testing, which fails to verify whether tests adequately capture behavioral relationships implied by APIs and documentation. The paper proposes the first evaluation framework based on inferred metamorphic relations (MRs): it automatically derives MRs using a UI-specific taxonomy from source code and documentation, aligns test executions to these MRs through deterministic and semantic analysis, and introduces relation-level MR coverage as a novel metric. By treating inferred MRs as empirical benchmarks for behavioral validation, the approach exposes verification gaps invisible to traditional coverage metrics—particularly in weak-oracle scenarios. Empirical results across three LLM configurations show MR coverage ranging only from 42.5% to 47.6%, substantially lower than MR reachability; uncovered MRs are predominantly of the weak-oracle type, demonstrating that MR coverage meaningfully complements conventional metrics and offers practical utility in fault detection and issue mapping.

behavioral validationmetamorphic relationstest coverage

This study addresses the lack of traceable, structured linkage between high-level requirements and low-level automated testing in AI-enabled cyber-physical systems, which hinders compliance with regulatory demands for verifiable evidence. To bridge this gap, the paper introduces VNVSpec, a novel framework that enables end-to-end automated traceability and closed-loop verification from high-level engineering requirements to test cases. VNVSpec employs machine-readable verification and validation (V&V) specifications to support requirement ingestion, quality checks, metric-driven decomposition, test result association, and generation of audit-ready reports, all integrated into a continuous integration pipeline. Empirical evaluation demonstrates that the approach verifies 36 requirements against 449 tests in linear time, scales to tens of thousands of artifacts, and is fully reproducible through open-sourced code, test suites, and benchmark scripts.

high-level requirementslow-level testsmachine-readable specifications