counterexample-guided abstraction refinement

Design, build, or analyze iterative abstraction–refinement workflows that produce candidate models or plans, detect counterexamples or conflicts in those candidates, and automatically add constraints to the abstraction to eliminate the observed counterexamples. Implement and tune the refinement loop and its encodings or solvers (for example using SMT-based encodings or conflict-based search hybrids such as SMT-CBS) so the process converges to a verified or conflict-free solution.

counterexample-guidedabstractionrefinement

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the common omission in current large language model evaluations of code generation—the iterative refinement process inherent in real-world programming and the models’ capacity for self-correction using feedback. The authors propose a novel framework that leverages execution-based feedback, such as compilation errors and test failures, to systematically investigate how reasoning and non-reasoning models utilize such signals across multiple programming languages. Through multidimensional categorization of code failures and extensive cross-model, cross-language experiments, they demonstrate that reasoning models consistently improve over iterations and significantly outperform non-reasoning counterparts. While syntactic and runtime errors prove relatively amenable to correction, logical and algorithmic errors remain challenging, thereby delineating the current limits of feedback-driven repair mechanisms.

code correctionexecution feedbackiterative refinement

This work addresses the prevailing lack of systematic understanding of foundational formal theories in current AI compiler design, which hinders rigorous evaluation of the completeness and desirability of intermediate representations and compilation abstractions. For the first time, it systematically establishes precise correspondences between core mechanisms of MLIR—such as term rewriting systems, refinement calculi, and abstract interpretation—and classical formal theories. By grounding compiler abstractions in formal semantics, the paper clarifies the theoretical underpinnings of these constructs, articulates a precise notion of “design completeness,” and provides assessable criteria and guiding principles to navigate trade-offs between engineering pragmatism and theoretical ideals.

abstraction designAI model compilationcompiler infrastructure

Counterexample-Guided Abstraction Refinement for Generalized Graph Transformation Systems (Full Version)

Apr 11, 2025
BK
Barbara König
🏛️ University of Duisburg-Essen | University of Twente

This paper addresses the (un)reachability verification problem between initial and error states in graph transformation systems specified by first-order nested conditions. Due to infinite state spaces, this problem is generally undecidable. To tackle it, we propose the first Counterexample-Guided Abstraction Refinement (CEGAR) framework tailored for general graph transformation systems with first-order nested condition constraints, integrating abstract interpretation, predicate abstraction, and graph transformation semantics into a terminating automated verification procedure. Our key contribution is a novel abstraction and refinement mechanism specifically designed for nested conditions, enabling precise unreachability proofs for complex, structured error states. We validate the effectiveness and practicality of our approach on multiple case studies. The method provides a new, generic pathway for formal verification of reactive systems governed by structural constraints.

Ensuring automated verification of nested conditionsHandling infinite state spaces via abstraction refinementVerifying graph reachability in transformation systems

Refinement-Types Driven Development: A study

Sep 18, 2025
FD
Facundo Domínguez
🏛️ Tweag

SMT solvers are traditionally confined to formal verification, limiting their utility in everyday programming tasks—particularly in enhancing standard type checkers’ capabilities for program composition and complex scoping (e.g., compiler binders). Method: We propose deep integration of refinement types into the compiler’s static checking pipeline, leveraging SMT solvers to automatically discharge refinement constraints. Building on Liquid Haskell, we design and implement an SMT encoding prototype supporting the theory of finite maps. Contribution/Results: Our approach significantly improves type-checking precision and developer experience by enabling richer behavioral specifications and more precise reasoning within the type system. Evaluation demonstrates substantial gains in correctness and constructibility for compiler binder scopes and other realistic scenarios. The resulting static assurance mechanism bridges practical usability with formal reliability, extending SMT-based reasoning beyond verification into mainstream compilation and development workflows.

Advocating broader SMT solver use beyond formal verificationEnhancing type checkers through SMT-integrated refinement typesSimplifying programming tasks with refinement types and solvers

Latest Papers

What's happening recently
View more

This work addresses the challenge of repeatedly verifying the safety of cyber-physical systems during iterative design, where frequent reconfigurations necessitate costly global revalidation. To overcome this, the paper introduces the principle of “refactoring-as-proposition,” which, for the first time, formalizes hybrid system refactoring as provable logical propositions. Leveraging differential refinement logic (dRL), the approach uniformly characterizes system properties and their preservation across refactorings—including those involving auxiliary variables—and enables localized verification. By transferring safety proofs from the original system to its refactored variant, the method substantially reduces verification complexity, facilitates automated or modular proof construction, and eliminates the need for exhaustive revalidation of the entire system.

cyber-physical systemshybrid systemsproperty preservation

This work addresses the optimal stopping problem in the self-refinement process of foundation models, aiming to determine the best termination point with minimal computational cost. The iterative refinement procedure is formulated as an optimal stopping problem that balances expected performance gains against computational overhead. The paper introduces, for the first time, an efficient and computable stopping strategy that dynamically decides when to halt refinement by integrating stochastic approximation, in-context learning, and external feedback mechanisms. Experimental results on code generation benchmarks demonstrate that the proposed method significantly reduces computational costs while maintaining or even surpassing the performance of existing approaches, thereby validating its effectiveness and practicality.

cost-efficiencyfoundation modelsiterative refinement

This work investigates whether model ensembles within the 1–3B parameter range can enhance code generation performance through execution feedback and pipeline architectures. We construct a generate-and-refine pipeline based on small language models, incorporate an execution feedback mechanism, and employ a NEAT-inspired evolutionary algorithm to search for optimal topologies. Our experiments reveal that execution feedback is pivotal—yielding performance gains exceeding four standard deviations on HumanEval and MBPP, primarily by correcting runtime errors—whereas increased topological complexity offers no significant benefit. The refinement component’s capability outweighs the identity of the generator, and single-run evaluations tend to overestimate evolutionary improvements; early stopping proves essential to prevent performance degradation. Moreover, specialized code models consistently outperform all combinations of general-purpose models.

code generationexecution feedbackmodel composition

Hot Scholars

TC

Taolue Chen

School of Computing and Mathematical Sciences, Birkbeck, University of London
Software EngineeringProgram Analysis and VerificationMachine learning
NP

Nir Piterman

Professor in Computer Science, University of Gothenburg and Chalmers University of Technolog, Sweden
VerificationAutomataLogicGames
RK

Roham Koohestani

Research Assistant, AISE TU Delft
AI4SEMachine LearningSoftware Engineering
MF

Marie Farrell

The University of Manchester
Formal MethodsAutonomous Robotic SystemsSoftware Verification
JW

Ji Wang

National University of Defense Technology
Distributed ComputingMachine Learning