Score
Design and build training pipelines and models that learn to score, filter, or rerank generated program or query candidates using their execution outcomes or symbolic checks as automatic supervision; this includes integrating verifiers into the generation loop, producing and sampling large candidate sets, and creating automated candidate-labeling procedures based on execution or symbolic verification. Work covers specifying verifier architectures and loss functions that consume execution-based signals, orchestrating verifier-guided training or inference, and evaluating how execution/symbolic supervision scales verifier performance without manual annotation.
This work addresses the limitations of process reward modeling (PRM)—namely, its heavy reliance on labor-intensive step-level human annotations and poor generalization. We propose a generative, long-chain, verifiable chain-of-thought (CoT) PRM paradigm. Our method fine-tunes large language models to generate fine-grained, verifiable long CoTs, integrating reward-guided search and best-of-N inference. Crucially, it achieves efficient training using only 1% of the process labels in PRM800K. The core contribution is the first end-to-end generative PRM framework, eliminating dependence on discriminative modeling and manual annotation. Experiments demonstrate state-of-the-art performance across ProcessBench, MATH-500, and AIME’24; cross-domain gains of +8.0% on GPQA-Diamond and +4.5% on LiveCodeBench; and a +7.2% improvement in verification accuracy under identical token budgets.
Large language models (LLMs) exhibit insufficient reasoning capabilities for complex programming tasks: process supervision relies on costly and error-prone reward modeling, while outcome supervision struggles to coordinate multi-step reasoning. To address this, we propose a novel “outcome-refinement-as-process” supervision paradigm that eliminates explicit reward modeling and instead leverages program execution feedback—such as runtime outputs and error traces—as label-free, reliable intermediate supervision signals. Our approach integrates tree-based multi-path exploration with a lightweight model adaptation framework to enable efficient, execution-guided reasoning. Evaluated across five LLMs and three benchmark datasets, our method achieves average improvements of 26.9% in code correctness and 42.2% in execution efficiency. Notably, it significantly boosts the performance of smaller models on algorithmic competition–style tasks. This work establishes a scalable, low-overhead paradigm for complex programming reasoning, grounded in direct execution feedback rather than surrogate reward signals.
Minor inaccuracies in process verifiers are readily amplified during language model generation, leading to catastrophic failures; moreover, training high-quality verifiers is prohibitively expensive. Method: We propose the test-time sampling algorithm VGB, the first to introduce the Sinclair–Jerrum random walk from theoretical computer science into language generation. VGB constructs a joint-probability-guided dynamic backtracking mechanism over the autoregressive generation tree, enabling robust responses to erroneous verifier signals. Contribution: We establish a falsifiable robustness-theoretic framework that formally uncovers the intrinsic connection between process verification and approximate sampling. Empirically, VGB significantly mitigates performance degradation induced by verifier errors across both synthetic and real-world tasks, consistently outperforming baselines on multiple evaluation metrics.
Existing process supervision data lacks controllability in error location, type, and trajectory consistency, hindering effective training of process reward models. This work proposes a controllable and verifiable synthetic framework that injects template-aware errors into intermediate steps of correct symbolic reasoning chains, recomputes subsequent steps via state propagation, and ensures the first error cannot be logically derived from prior context, thereby generating natural language process pairs that are internally consistent and precisely localize the first mistake. For the first time, this approach enables fine-grained control over error position, type, and trajectory coherence while incorporating a verification mechanism to guarantee supervision reliability. Experiments demonstrate that the synthesized data significantly improves Best-of-8 reranking performance on logical reasoning tasks and generalizes effectively to mathematical reasoning, further revealing that pinpointing the first error is substantially more challenging than classifying entire step sequences—highlighting the necessity of fine-grained process supervision.
This work addresses the challenge that large language models struggle to effectively detect property violations in program verification, particularly exhibiting significant performance degradation on long programs. The authors propose a novel approach that leverages error traces generated by the symbolic execution engine Soteria as continued pretraining data for the Qwen3-8B model, combined with chain-of-thought reasoning at inference time to enhance semantic understanding of programs. Using only approximately 3,000 such traces, this method improves violation detection accuracy by over 17 percentage points, enabling the 8B-parameter model to surpass a 32B-parameter counterpart without chain-of-thought reasoning. The approach demonstrates balanced performance across five SV-COMP property categories and generalizes to unseen property types, validating the superadditive effect arising from the synergy among error trace semantics, formatting, and chain-of-thought reasoning.
This work addresses the bottleneck of loop invariant synthesis in formal verification by introducing VerIbmc, the first fully local neurosymbolic framework that operates without reliance on cloud-based large language model APIs. The approach integrates deterministic symbolic reasoning, locally deployed open-source large language models (ranging from 7B to 120B parameters), the ESBMC model checker, and a structured feedback mechanism, supporting both Chain-of-Thought and Tree-of-Thought prompting strategies. Evaluated on 499 benchmark problems, the best configuration (GPT-OSS-120B) solves 431 instances (86.4%), matching the performance of state-of-the-art cloud-based tools. Notably, the symbolic component alone solves 75 problems and substantially enhances the efficacy of weaker models, achieving efficient verification while preserving code privacy and minimizing computational cost.
Existing query-first data synthesis approaches struggle to generate valid and executable tool-use sequences. This work proposes SyntheticAgentTraceQA, a novel framework that introduces an "execution-first" paradigm: it first constructs high-level workflows, maps and validates feasible tool trajectories, and then synthesizes corresponding natural language tasks and reference answers. The method integrates dependency-aware tool assignment, trajectory validation in a controlled environment, and reasoning-augmented annotation generation, followed by fine-tuning and evaluation using the Qwen model. Experimental results demonstrate that this framework substantially improves large language model (LLM) agents’ tool execution accuracy, trajectory consistency, and answer quality. Furthermore, the study reveals that masked supervision outperforms full supervision for models at the 9B scale.