Score
Designs and implements post-hoc answer verification systems and routing policies that score candidate results with per-instance (e.g., per-cell) metrics and emit verdicts or confidence signals to decide automated acceptance or human-in-the-loop escalation. This work includes building answer verifiers and routing logic and defining differential scoring and reward functions (for example top-k Pearson or RMSE proximity, Spearman relationships, or other per-cell metric definitions) to align outputs with reference patterns or expected activity.
Large language models (LLMs) have long relied on outcome-based reward modeling—evaluating only final answers—leading to insufficient interpretability and robustness in reasoning processes. Method: This paper systematically introduces the Process Reward Modeling (PRM) paradigm, shifting supervision from outcome-level to step- or trajectory-level reasoning evaluation. We establish a comprehensive methodology encompassing data construction, fine-grained reward modeling, test-time scaling, and RLHF integration. Contribution/Results: We empirically validate PRM across diverse domains—including mathematics, code generation, natural language, multimodal reasoning, and robotic agent tasks. To our knowledge, this is the first work to characterize the design space and core challenges of PRMs across multiple domains, releasing a cross-task benchmark, practical implementation guidelines, and open-source resources. Our framework provides foundational theoretical insights, actionable technical pathways, and empirical support for trustworthy reasoning alignment in LLMs.
This work addresses the limitations of existing open-domain post-training methods, which rely on scalar rewards assigned after generation and struggle to explicitly model prompt-specific local requirements, holistic preferences, and hard constraints. The authors propose a prompt-level reward specification framework that decouples reward definition from computation: task-adaptive scoring rubrics and executable constraint checkers are constructed offline based on prompts prior to training, and combined with global quality scores to produce a hybrid reward signal. This framework enables explicit reward specification without human preference labels, reference answers, or a separate reward model—supporting both offline and online reinforcement learning. Experiments demonstrate significant improvements in response ranking across multiple open-domain benchmarks and effective online policy optimization, while ablation studies confirm the complementary roles of individual components.
This work addresses a critical limitation in existing answer verifiers, which assess only final answer correctness while ignoring errors in the reasoning process—thereby misclassifying correctly answered but erroneously derived solutions as valid. To remedy this, the paper introduces PRIME, a novel benchmark that pioneers the concept of process–result alignment verification, systematically evaluating verifiers’ ability to jointly judge reasoning consistency and answer correctness on challenging STEM tasks in mathematics and engineering. Leveraging a high-quality dataset curated via consistency-based filtering, the authors propose RLVR, a training paradigm integrating process-aware reinforcement learning with a verifiability-aware reward mechanism. Experiments demonstrate that RLVR-enhanced verifiers achieve performance gains of 8.29%, 9.12%, and 7.31% on AIME24, AIME25, and Beyond-AIME, respectively, with verifier accuracy strongly correlated to RLVR efficacy (R² > 0.92).
This work proposes a Verifiable Process Reward Model (VPRM) to address the limitations of existing process supervision methods, which rely on neural discriminators to evaluate intermediate reasoning steps of large language models and are thus prone to opacity, bias, and reward hacking, often failing to enforce domain-specific reasoning rules. VPRM introduces, for the first time, a deterministic, rule-based programmatic verifier into a reinforcement learning framework, enabling interpretable, auditable, and deception-resistant supervision of intermediate reasoning. Evaluated on the task of bias risk assessment in medical evidence synthesis, VPRM substantially improves logical consistency and evidential grounding of model reasoning, achieving up to a 20% absolute gain in F1 score across multiple datasets and outperforming outcome-only verification baselines by 6.5%.
Synthetic verification methods—such as self-generated test cases and reward modeling—lack systematic evaluation for code correctness assessment. Method: We introduce four novel benchmarks—HE-R, HE-R+, MBPP-R, and MBPP-R+—unifying programming evaluation datasets into dual-modality tasks: scoring and ranking. We propose the first standardized evaluation framework tailored to synthetic verification capability, enabling multi-dimensional accuracy assessment and attribution analysis. Contribution/Results: Experiments reveal a synergistic gain between reasoning model scale and test case quantity; mainstream LLMs significantly improve test generation quality; and increasing test cases consistently enhances verification accuracy. Cross-model evaluation across multiple LLMs identifies key determinants of verification performance, including test diversity, oracle reliability, and model calibration. This work establishes foundational infrastructure and empirical insights for advancing trustworthy code verification via synthetic methods.
This work addresses the challenge in reinforcement learning where reliance solely on sparse end-task rewards impedes effective optimization of reasoning trajectories, due to the absence of intermediate feedback before task success and the inability to distinguish efficient from redundant paths afterward. To overcome this, the authors propose SCOPE-RL, a two-stage framework that introduces verifiable prefix-decomposition rewards via answer-hidden subproblem chains prior to success, and subsequently employs correctness-gated process-shaped rewards to refine trajectories while preserving the GRPO update mechanism. The approach innovatively integrates pre- and post-success process-level reward signals through adaptive scaffolding and quality-aware process reinforcement learning, complemented by a Step-Quality Evaluation Protocol validated by human experts. Evaluated on Qwen3-8B-Instruct, SCOPE-RL achieves up to an 11.2 percentage point gain in accuracy and reduces reasoning tokens by 27.1%, demonstrating consistent generalization across diverse models and algorithms.
This work addresses the lack of a reproducible evaluation framework for meta-decision strategies—such as task decomposition and tool invocation—in existing agent systems. We introduce MetaRoute-Bench, the first open benchmark enabling fine-grained analysis of meta-decision routing, comprising 180 synthetic tasks, 8 distinct strategies, and 30 random seeds per configuration. Evaluation employs offline seeded execution and multidimensional metrics—including success rate, cost, and latency—to ensure fair comparison. Experiments demonstrate that task-aware compositional strategies achieve a significantly higher success rate (79.4%) compared to static strategies (76.7%), single-step routing (67.4%), and direct answering (52.9%), with only marginal increases in cost (4.7%) and latency (6.4%). Ablation studies further confirm the critical contributions of compositional operations and verification mechanisms. Code and execution trajectories are publicly released.
This work addresses the challenge of "reasoning collapse"—a phenomenon wherein large language models, when applied to subjective tasks such as content moderation, over-reason and fail to align with human preferences. The study reveals a strong correlation between reasoning style and verification efficacy, and for the first time formally identifies and characterizes reasoning collapse. To mitigate this issue, the authors propose a dynamic persona-aware reasoning routing mechanism, integrated with conditional length penalty fine-tuning, reinforcement learning with verifiable rewards (RLVR), and large-scale role simulation. Experiments across four real-world recommendation platforms demonstrate that the proposed approach effectively suppresses reasoning collapse, achieving up to a 0.38 improvement in macro-F1 solely through adjustments to reasoning persona style, thereby significantly enhancing alignment performance on subjective tasks.