Score
Design and build methods and pipelines that generate task specifications from desired or observed outcomes and simultaneously produce executable verification artifacts (tests, environment checks, or simulators) that demonstrate the synthesized tasks actually yield those outcomes. Implement and analyze backward/inverse task-synthesis algorithms and test-driven, environment-based verification frameworks that guarantee task labels or success criteria are correct by construction.
The test-time scaling (TTS) field lacks a systematic survey, unified taxonomy, and principled analysis of verification methods. Method: We propose the first taxonomy framework for TTS verifiers, structured along three dimensions—verifier type, training paradigm (prompt-based guidance, discriminative/generative fine-tuning), and application mode—and integrate search-space exploration with candidate-output scoring for efficient inference optimization. We further provide a comprehensive survey of existing verification techniques and release an open-source verification resource repository. Contribution/Results: Our framework fills a critical gap in systematic TTS verification research, significantly improving the accuracy and reliability of large language model inference. It establishes a standardized foundation and reproducible benchmark for verifier design in TTS, enabling principled development and evaluation of verification mechanisms.
To address the challenges of simultaneously generating semantically consistent yet stylistically diverse multi-artifact programming exercises—namely source code, test specifications, and natural language descriptions—this paper proposes a compositional generation framework grounded in abstract syntax building blocks. The framework defines reusable syntactic abstractions and integrates templated mapping with multi-objective instantiation to ensure intent preservation and cross-modal co-generation. Its key innovations include: (i) enabling style-controllable, diverse outputs while guaranteeing semantic consistency; and (ii) providing a highly configurable generation interface that substantially reduces customization effort for new tasks. Experimental evaluation demonstrates that the approach outperforms existing baselines across three critical dimensions: generation quality, output diversity, and system extensibility.
This work addresses the significant disparity in verifiability among semantically equivalent yet structurally diverse programs, a key bottleneck in generating high-assurance software. The authors propose Diversify2Verify, a novel approach that leverages large language models to synthesize diverse recursive and imperative implementations of the same task, integrates the Why3 platform for automatic contract inference and formal verification, and introduces a verifier-guided annotation repair mechanism to enhance verifiability. This study is the first to systematically expose the verifiability gap across equivalent program variants and establishes a new paradigm wherein implementation diversity drives improved verification success. Evaluated on a benchmark of 73 tasks, the method yields 154 verifiable programs after two rounds of repair, with at least one successfully verified variant for 67.1% of the tasks—substantially outperforming baseline approaches.
This work investigates the role of task decomposition in program synthesis, examining its impact on generalization and subgoal validity by comparing the explicit decomposition framework ExeDec against the decomposition-free, execution-driven framework REGISM. Method: We propose a novel synthesis paradigm that jointly optimizes iterative execution feedback, subgoal modeling, and code generation, and introduce a cross-task generalization evaluation framework. Contribution/Results: (1) Execution-driven learning alone serves as a critical performance driver; (2) explicit decomposition substantially improves length generalization and compositional concept learning; (3) despite lacking explicit decomposition, REGISM matches or surpasses ExeDec across multiple metrics, and its implicit decomposition aligns more closely with human-annotated subgoal structures. Our study is the first to reveal the fundamental trade-off—decomposition is not strictly necessary but can yield measurable gains—thereby offering a new perspective on task structure modeling in program synthesis.
Automatically generating high-quality, correct, and complete formal specifications—such as those in JML—remains a significant challenge: existing approaches often produce specifications that pass syntactic validation yet suffer from semantic inaccuracies or insufficient coverage. This work proposes VeriAct, a novel framework that introduces Spec-Harness, the first evaluation mechanism capable of precisely assessing both correctness and completeness of generated specifications. VeriAct further establishes the first verification-guided agent system, leveraging large language models within a closed-loop iterative process that integrates code execution, formal verification, and feedback signals to collaboratively synthesize and repair specifications. Experimental results demonstrate that VeriAct substantially outperforms current methods on two benchmarks, yielding specifications that not only satisfy verifiers but also achieve higher standards of semantic correctness and completeness.
Current LLM-native software engineering lacks a systematic practical framework—particularly in verification and falsification—necessitating unified task taxonomies and prompt-pattern conceptualizations. Method: We conduct a systematic literature review of over 100 papers, employing bibliometric analysis and conceptual clustering to map, classify, and abstract LLM-based downstream tasks in software engineering (SE). Contribution/Results: We propose the first fine-grained SE-specific taxonomy for LLM downstream tasks, encompassing six core clusters: testing, fuzzing, bug localization, vulnerability detection, static analysis, and program verification. Our taxonomy uniquely balances cross-task abstraction with task-specific variation modeling, uncovering generalizable prompt-engineering principles. It provides a foundational framework for targeted LLM adaptation, benchmark construction, and empirically grounded engineering practice in SE.
High-quality training data for terminal-based intelligent agents remains scarce, as existing synthetic methods often yield ambiguous instructions, shallow execution paths, and fragile tests that fail to provide effective learning signals. To address this, this work proposes a structured, high-fidelity task synthesis paradigm: it samples task candidates based on a multidimensional capability taxonomy, constructs them in depth through evidence-guided grounding in real technical documentation, and ensures executability and challenge via Dockerized environment instantiation, scoring-gated test generation, and strict Fail-to-Pass validation. The resulting CLI-Universe-6K dataset comprises 6,000 verifiable agent trajectories. Fine-tuning Qwen3-32B on this dataset achieves a 33.4% success rate on Terminal-Bench 2.0, establishing a new state-of-the-art among open-source models of 32B parameters or fewer and surpassing several larger models.
This work addresses the scarcity of high-quality, verifiable, and diverse task data that hinders large-scale training of terminal-based intelligent agents. Existing synthetic approaches often suffer from a disconnect between task generation and execution and rely heavily on pre-existing repositories, limiting diversity and scalability. To overcome these limitations, the authors propose modeling the task synthesis process itself as a Terminal-Bench–formatted terminal task, enabling closed-loop iterative generation, execution, and validation within real containerized environments. Their method enhances diversity and realism through multi-stage task specification, decoupling of task dimensions, and augmentation with external materials, while employing an LLM-as-Judge mechanism for quality filtering. Using only 3,221 synthesized trajectories for fine-tuning, Qwen3-14B and Qwen3-32B achieve Avg Pass@1 scores of 22.5% and 31.8%, respectively, on Terminal-Bench 2.0—significantly outperforming concurrent methods with substantially less training data.
Existing query-first data synthesis approaches struggle to generate valid and executable tool-use sequences. This work proposes SyntheticAgentTraceQA, a novel framework that introduces an "execution-first" paradigm: it first constructs high-level workflows, maps and validates feasible tool trajectories, and then synthesizes corresponding natural language tasks and reference answers. The method integrates dependency-aware tool assignment, trajectory validation in a controlled environment, and reasoning-augmented annotation generation, followed by fine-tuning and evaluation using the Qwen model. Experimental results demonstrate that this framework substantially improves large language model (LLM) agents’ tool execution accuracy, trajectory consistency, and answer quality. Furthermore, the study reveals that masked supervision outperforms full supervision for models at the 9B scale.
This work addresses the high cost of manually writing formal specifications and the limitations of existing large language model (LLM)-based approaches that require white-box access to source code, thereby posing intellectual property and deployment constraints. The authors propose a black-box-driven method that leverages only test code and dynamic execution traces to generate candidate Java Modeling Language (JML) specifications via an LLM. These candidates are locally validated using bounded model checking, and an iterative feedback loop refines them based on verification outcomes. This approach is the first to enable fully automated formal specification generation without any access to the program’s internal structure. Evaluated on the SpecGenBench benchmark, it demonstrates that test-derived information effectively guides specification synthesis, while also highlighting critical challenges in checker compatibility and diagnostic feedback, substantially enhancing industrial applicability.