Score
Designs and constructs explicit problem instances, constructive countermodels, or input examples that demonstrate algorithmic or theoretical failure modes, violate formal invariants, or realize lower-bound behavior. Searches for and analyzes invariant-violating inputs, evaluates counterexample coverage, and uses these concrete constructions to isolate mechanisms of failure and to refine proofs, bounds, or algorithmic remedies.
This work addresses the challenge that counterexamples generated by formal verification often consist of numerous low-level Boolean variables, rendering them difficult for developers to interpret at the application-domain level. To bridge this gap, the paper proposes a novel hierarchical explanation method that integrates predicate relevance metrics with dependency graph analysis—a first-time fusion of these two techniques—to automatically extract human-readable, domain-oriented explanations from logical formulas. By leveraging formal modeling and a dedicated explanation-generation algorithm, the approach produces concise and semantically clear descriptions of failure causes across multiple case studies. Empirical results demonstrate that the method significantly outperforms existing techniques, offering effective support for fault localization in practical verification tasks.
This work proposes a novel paradigm that bridges the long-standing divide between testing and formal verification in traditional software validation, enabling them to synergistically enhance both efficiency and quality. Grounded in Design by Contract, the approach leverages the counterexample generation capability of SMT solvers to transform formal verification tools into an integrated engine for automated testing and repair. Within a unified framework, the method simultaneously achieves three key objectives: automatic generation of test cases for faulty programs, construction of regression test suites with full coverage for correct programs, and correctness-guaranteed program repair. This represents the first integration of verification, testing, and repair into a single cohesive methodology.
Current mathematical large language models (LLMs) rely heavily on proof examples in training data and lack deep conceptual understanding of theorems’ underlying principles. Method: We propose a novel counterexample-driven conceptual reasoning paradigm to overcome this mathematical reasoning bottleneck. Contribution/Results: (1) We introduce CounterMATH—the first university-level benchmark explicitly designed for counterexample generation and conceptual discrimination; (2) we develop a scalable, prompt-driven automated data engineering framework for fine-grained counterexample synthesis and training data curation; (3) through multi-model comparative evaluation and attribution analysis, we empirically expose systematic deficiencies of mainstream mathematical LLMs in counterexample-based reasoning, and demonstrate that targeted fine-tuning significantly improves both conceptual comprehension and formal proof generation capabilities. This work establishes a new standard and actionable pathway for evaluating and enhancing mathematical reasoning in LLMs.
This work addresses the challenge of effectively improving large language models when confronted with heterogeneous, domain-specific, and hard-to-control feedback. The authors propose a counterexample-guided learning framework that leverages a formal verifier to generate precise counterexamples, iteratively refining the model’s generation of regular expressions. Key innovations include a novel counterexample-guided refinement strategy—incorporating regularization and symbolic counterexample clustering—and a multi-agent reflective repair loop. Experimental results demonstrate substantial performance gains: on the most challenging tasks, success rates improve from 3.2% to 38.1%, and in another domain, from 38.9% to 74.1%, significantly enhancing sample efficiency and the ability to learn complex expressions.
This paper addresses the (un)reachability verification problem between initial and error states in graph transformation systems specified by first-order nested conditions. Due to infinite state spaces, this problem is generally undecidable. To tackle it, we propose the first Counterexample-Guided Abstraction Refinement (CEGAR) framework tailored for general graph transformation systems with first-order nested condition constraints, integrating abstract interpretation, predicate abstraction, and graph transformation semantics into a terminating automated verification procedure. Our key contribution is a novel abstraction and refinement mechanism specifically designed for nested conditions, enabling precise unreachability proofs for complex, structured error states. We validate the effectiveness and practicality of our approach on multiple case studies. The method provides a new, generic pathway for formal verification of reactive systems governed by structural constraints.
This study addresses the challenge computer science students face in transitioning from operational to structural thinking and in mastering the construction of loop invariants within formal methods. To this end, it proposes a Graphical Loop Invariant-Based Programming (GLIBP) approach that guides learners to first model problems using graphical loop invariants (GLIs) via block diagrams and then derive code from these representations. The work introduces, for the first time, an integrated consistency analysis and automated feedback mechanism that jointly evaluates GLIs and their corresponding code. The accompanying educational tool not only provides personalized guidance to students but also assists instructors in designing programming assignments, thereby significantly enhancing students’ comprehension and application of formal methods.
This work addresses the limited interpretability of Computation Tree Logic (CTL) model checking results, which stems from the absence of intuitive, visualizable evidence forms—CTL counterexamples being notably harder to comprehend than Linear Temporal Logic (LTL) traces. The paper proposes a unified evidence framework for CTL over explicit-state models, capable of representing both witnesses for satisfied properties and counterexamples for violations. Minimal evidence structures are formally defined for each temporal operator, and a human-centric visualization scheme is developed by integrating formal reasoning with graphical techniques. A complete toolchain implementing this approach is presented. All theoretical claims are rigorously proven, significantly enhancing the readability and explainability of CTL model checking outcomes.
This work addresses critical reliability issues in existing Lean theorem-proving benchmarks, where inconsistencies between formal statements and informal problem descriptions, along with susceptibility to trivial or adversarial solutions, undermine evaluation validity. To tackle this, the study introduces the first fault taxonomy for formal mathematical datasets and develops an automated auditing toolkit integrating static program analysis, formal verification, semantic auditing via prompt engineering, and manual validation. Applying this framework, the authors conduct a large-scale audit of five prominent benchmarks, uncovering 4,833 issues—including 398 severe defects—and demonstrate that uncorrected flaws significantly distort prover rankings. The paper releases both the auditing tools and corrected dataset snapshots to foster reproducible and trustworthy evaluations in theorem proving.
This work addresses a critical yet overlooked reliability issue in code generated by large language models (LLMs): despite passing compilation and unit tests, such code often fails in deployment due to structural inconsistencies—such as missing configurations, invalid imports, or omitted security controls—that evade detection by conventional CI/SAST tools. The paper introduces the “patchwork problem” to characterize these cross-module global defects, proposes an eight-category taxonomy specific to LLM-generated code, and formalizes structural consistency via invariants derived from a multidimensional code graph encompassing imports, calls, dependencies, configurations, and routing. Building on this foundation, the authors design a hybrid verification framework that integrates traditional static analysis with custom graph-based invariant checkers to precisely identify structural flaws invisible to existing tools. Empirical evaluation reveals that such defects are pervasive across major LLMs under diverse prompting strategies and exhibit distinct model-specific patterns.