Score
Designs and implements systems that generate executable programs by producing, storing, and recombining independently synthesized modules—learning multiple implementations per module and composing them into full-program candidates. Includes methods to organize, search, and rank module implementations so as to generate many candidate programs efficiently and to shift expensive inference work toward cheaper verification or evaluation steps.
This work investigates the role of task decomposition in program synthesis, examining its impact on generalization and subgoal validity by comparing the explicit decomposition framework ExeDec against the decomposition-free, execution-driven framework REGISM. Method: We propose a novel synthesis paradigm that jointly optimizes iterative execution feedback, subgoal modeling, and code generation, and introduce a cross-task generalization evaluation framework. Contribution/Results: (1) Execution-driven learning alone serves as a critical performance driver; (2) explicit decomposition substantially improves length generalization and compositional concept learning; (3) despite lacking explicit decomposition, REGISM matches or surpasses ExeDec across multiple metrics, and its implicit decomposition aligns more closely with human-annotated subgoal structures. Our study is the first to reveal the fundamental trade-off—decomposition is not strictly necessary but can yield measurable gains—thereby offering a new perspective on task structure modeling in program synthesis.
To address the challenges of simultaneously generating semantically consistent yet stylistically diverse multi-artifact programming exercises—namely source code, test specifications, and natural language descriptions—this paper proposes a compositional generation framework grounded in abstract syntax building blocks. The framework defines reusable syntactic abstractions and integrates templated mapping with multi-objective instantiation to ensure intent preservation and cross-modal co-generation. Its key innovations include: (i) enabling style-controllable, diverse outputs while guaranteeing semantic consistency; and (ii) providing a highly configurable generation interface that substantially reduces customization effort for new tasks. Experimental evaluation demonstrates that the approach outperforms existing baselines across three critical dimensions: generation quality, output diversity, and system extensibility.
Program synthesis tools suffer from poor reusability, high integration costs, and limited extensibility. To address these challenges, this paper introduces SynthLib—a modular, Julia-based program synthesis library that decouples grammar specification, problem modeling, synthesis algorithms, and benchmarking into composable, interoperable components. Its core innovation lies in a unified abstract interface and a lightweight architectural design, which drastically lowers the barrier to implementing new synthesis methods: reproducing classic synthesizers requires only dozens of lines of code, and integrating novel algorithms reduces average development time by 80%. SynthLib is rigorously validated on standard benchmarks—including SyGuS-Comp—demonstrating correctness, efficiency, and scalability. By providing a reusable, experiment-friendly foundation, SynthLib advances programmable synthesis research and facilitates rapid prototyping, comparative evaluation, and collaborative tool development.
Formal program specifications are notoriously difficult, error-prone, and inefficient to write manually. To address this, we propose a two-stage LLM-driven approach: dialogue-guided specification synthesis followed by mutation-based verification. First, multi-turn dialogues model complex semantic requirements; second, four mutation operators—insertion, replacement, deletion, and reordering—enable verifiability-driven selection, eliminating reliance on rigid templates or syntactic grammars. Our method integrates code understanding, prompt engineering, and heuristic verifiability assessment. Evaluated on SV-COMP and a custom Java benchmark comprising 385 programs, it generates 279 verifiable specifications. These achieve significantly higher completeness and accuracy than pure-LLM baselines and classical tools (e.g., Houdini, Daikon). To our knowledge, this is the first approach to achieve both high coverage and formal verifiability in fully automated specification generation.
Program synthesis suffers from exponential search-space explosion and low efficiency under weak prior knowledge. Method: This paper proposes a “Decompose–Align–Solve” framework: (1) interpretable decomposition of input-output structures based on semantic units; (2) incorporation of Structure-Mapping Theory (SMT) from cognitive science to establish cross-modal structural correspondences, guiding collaborative subprogram synthesis; and (3) realization of a decomposition-driven, scalable synthesis paradigm. Contribution/Results: This work is the first to systematically integrate SMT into program synthesis, enabling compositional program construction and explicit input-output structural alignment. Theoretical analysis shows time complexity reduced to O(m) and monotonic improvement in prediction accuracy with increasing samples. Experiments demonstrate significant gains over ILP baselines on string transformation tasks and, for the first time, achieve effective program synthesis on the ARC visual reasoning benchmark in scenarios where ILP is infeasible.
This work addresses the challenge that large language model agents face when programming from scratch, where entangled document comprehension, behavioral exploration, and code generation often lead to intent drift and error propagation. To mitigate this, the authors propose SpecFirst, a two-stage framework: in the first stage, a dedicated specification agent synthesizes structured behavioral specifications by integrating binary probing with documentation analysis; in the second stage, a code synthesis agent generates programs strictly adhering to these specifications. By explicitly decoupling specification construction from code generation—drawing inspiration from requirements engineering—this approach introduces behavioral specification extraction as an independent, primary step in agent-driven programming. Evaluated on all 200 instances of ProgramBench, SpecFirst improves test pass rates by 6.9%–21.3%, increases binary exploration coverage by 9.4%–18.5%, and yields earlier and more stable code construction.
It remains unclear whether current large language models genuinely understand program semantics in code generation, and there is a lack of systematic evaluation of their ability to generate executable behavioral specifications. This work proposes CodeSpecBench, the first benchmark supporting multi-granularity tasks at both function and repository levels, which expresses preconditions and postconditions as executable Python functions and emphasizes the correctness and completeness of specifications. Using an execution-driven protocol, it evaluates a model’s capacity to accept valid behaviors and reject invalid ones. Experiments across 15 state-of-the-art models reveal that the highest pass rate on repository-level tasks is only 20.2%, significantly lagging behind general code generation performance, thereby demonstrating that specification generation is substantially more challenging and that strong code generation capabilities do not equate to deep semantic understanding.
This work proposes a novel approach to program analysis and optimization leveraging large language models (LLMs). Addressing the challenge of effectively integrating source code and intermediate representation (IR) information—a limitation in existing methods—it introduces LLMCompiler, pre-trained on IR, and employs a chunked embedding and aggregation strategy to produce unified program-level embeddings. By innovatively unifying the semantics of source code and IR, the method achieves a 1.54% error rate on algorithm classification, representing a 12% improvement over the current state of the art. It also attains competitive accuracy in heterogeneous device mapping, significantly advancing the application of LLMs in program understanding and optimization.