Score
Designs, builds, and analyzes iterative systems that generate, validate, and refine structured concepts or formal specifications by producing candidate proposals (including from language models) and executing programmatic checks or tests against them. Implements hierarchical or search-based refinement and selection mechanisms that rank candidates by measurable performance or correctness metrics to navigate combinatorial or algebraic design spaces.
This study addresses the common omission in current large language model evaluations of code generation—the iterative refinement process inherent in real-world programming and the models’ capacity for self-correction using feedback. The authors propose a novel framework that leverages execution-based feedback, such as compilation errors and test failures, to systematically investigate how reasoning and non-reasoning models utilize such signals across multiple programming languages. Through multidimensional categorization of code failures and extensive cross-model, cross-language experiments, they demonstrate that reasoning models consistently improve over iterations and significantly outperform non-reasoning counterparts. While syntactic and runtime errors prove relatively amenable to correction, logical and algorithmic errors remain challenging, thereby delineating the current limits of feedback-driven repair mechanisms.
This work addresses the prevailing lack of systematic understanding of foundational formal theories in current AI compiler design, which hinders rigorous evaluation of the completeness and desirability of intermediate representations and compilation abstractions. For the first time, it systematically establishes precise correspondences between core mechanisms of MLIR—such as term rewriting systems, refinement calculi, and abstract interpretation—and classical formal theories. By grounding compiler abstractions in formal semantics, the paper clarifies the theoretical underpinnings of these constructs, articulates a precise notion of “design completeness,” and provides assessable criteria and guiding principles to navigate trade-offs between engineering pragmatism and theoretical ideals.
This work addresses the challenge of ensuring program correctness in natural language-to-code generation, which is often hindered by the absence of high-quality formal specifications. The authors propose VeriSpecGen, a framework that decomposes natural language requirements into atomic clauses through a traceable refinement mechanism, generates requirement-driven tests with explicit traceability mappings, and synthesizes formal specifications aligned with user intent by localizing and repairing faulty clauses upon verification failure. Integrating large language models (e.g., Claude Opus 4.5) with the Lean proof assistant, the approach leverages refinement trajectories to generate 343K training samples, substantially enhancing model generalization and reasoning capabilities. Evaluated on the VERINA SpecGen benchmark, VeriSpecGen achieves an accuracy of 86.6%, outperforming the best baseline by up to 31.8 percentage points and demonstrating a relative improvement of 62–106% in specification synthesis performance.
This paper addresses the narrow design-space exploration and severe information overload in LLM-assisted programming by proposing an IDE framework deeply integrated into the design process. Methodologically, it elevates LLMs from single-point code generators to collaborative design-space exploration agents—a novel conceptual shift—explicitly modeling problem variants, solution pathways, and implicit assumptions via a design decision graph, and enabling multi-perspective problem reformulation and parallel solution generation. Contributions include: (1) establishing the first LLM-IDE co-design paradigm specifically for program design-space exploration; (2) enabling design decisions to be traceable, comparable, and evolvable; and (3) demonstrating through user studies a statistically significant increase in design breadth. The work further identifies attention management as the core bottleneck in LLM-augmented IDEs, thereby providing a new benchmark for human-AI collaborative design research.
Formal program specifications are notoriously difficult, error-prone, and inefficient to write manually. To address this, we propose a two-stage LLM-driven approach: dialogue-guided specification synthesis followed by mutation-based verification. First, multi-turn dialogues model complex semantic requirements; second, four mutation operators—insertion, replacement, deletion, and reordering—enable verifiability-driven selection, eliminating reliance on rigid templates or syntactic grammars. Our method integrates code understanding, prompt engineering, and heuristic verifiability assessment. Evaluated on SV-COMP and a custom Java benchmark comprising 385 programs, it generates 279 verifiable specifications. These achieve significantly higher completeness and accuracy than pure-LLM baselines and classical tools (e.g., Houdini, Daikon). To our knowledge, this is the first approach to achieve both high coverage and formal verifiability in fully automated specification generation.
This work addresses the limitations of existing large language model (LLM)-driven heuristic design methods in combinatorial optimization, which often rely on manual trial-and-error or domain-specific knowledge and lack a systematic mechanism for improvement. To overcome this, the authors propose a structured framework that formalizes heuristic discovery as a language-guided program optimization process, comprising three modular phases: forward evaluation, backward feedback, and program update. This design enables an iterative and composable optimization workflow, unifying and generalizing prior approaches while allowing flexible enhancements through modularity. Empirical evaluation across four real-world combinatorial optimization tasks demonstrates that the proposed method significantly outperforms baseline techniques, achieving up to a 0.17 improvement in the QYI metric on unseen test instances.
This work addresses the high cost of manually writing formal specifications and the limitations of existing large language model (LLM)-based approaches that require white-box access to source code, thereby posing intellectual property and deployment constraints. The authors propose a black-box-driven method that leverages only test code and dynamic execution traces to generate candidate Java Modeling Language (JML) specifications via an LLM. These candidates are locally validated using bounded model checking, and an iterative feedback loop refines them based on verification outcomes. This approach is the first to enable fully automated formal specification generation without any access to the program’s internal structure. Evaluated on the SpecGenBench benchmark, it demonstrates that test-derived information effectively guides specification synthesis, while also highlighting critical challenges in checker compatibility and diagnostic feedback, substantially enhancing industrial applicability.
This work investigates whether model ensembles within the 1–3B parameter range can enhance code generation performance through execution feedback and pipeline architectures. We construct a generate-and-refine pipeline based on small language models, incorporate an execution feedback mechanism, and employ a NEAT-inspired evolutionary algorithm to search for optimal topologies. Our experiments reveal that execution feedback is pivotal—yielding performance gains exceeding four standard deviations on HumanEval and MBPP, primarily by correcting runtime errors—whereas increased topological complexity offers no significant benefit. The refinement component’s capability outweighs the identity of the generator, and single-run evaluations tend to overestimate evolutionary improvements; early stopping proves essential to prevent performance degradation. Moreover, specialized code models consistently outperform all combinations of general-purpose models.
Automated generation of verifiable formal specifications is often hindered by syntactic errors, logical inaccuracies, inadequate handling of control-flow structures, and the absence of dynamic error-correction mechanisms. This work proposes AutoReSpec, a novel framework featuring a two-stage collaborative generation mechanism that synergistically combines open- and closed-source large language models. By dynamically selecting prompting strategies based on program structure and invoking a collaborative model upon primary model failure, AutoReSpec leverages feedback from a formal verifier to iteratively refine specifications. Through structure-aware scheduling and a verification-in-the-loop architecture, the approach significantly enhances robustness and efficiency, achieving a 58.2% success rate and 69.2% completeness across 72 Java benchmarks, while reducing average evaluation time by 26.89% compared to existing methods.
This work presents the first systematic investigation into the capability of large language models (LLMs) to generate program specifications involving higher-order logical constructs, which are essential for expressing complex verification properties yet remain beyond the reach of existing LLMs that predominantly handle basic syntactic forms. The authors design four syntactic configurations spanning different levels of abstraction and establish a comprehensive evaluation framework to assess a range of representative LLMs on standard verification benchmarks. Experimental results demonstrate that LLMs can effectively produce valid higher-order logical expressions; moreover, integrating logical constructs with base syntax significantly enhances verification efficacy and robustness without substantially increasing verification overhead. The study also reveals distinct advantages of two refinement paradigms in specification generation.