Score
Design and implement algorithms, models, and evaluation procedures that attribute outcomes or rewards to individual actions, steps, or components of a system by learning how to allocate fractional or partial credit; this includes methods to identify weakest links, penalize unnecessary edits, and derive partial scoring rubrics. Build scoring and feedback mechanisms that aggregate multiple signals (including multimodal inputs) into composite scores and generate corrective feedback tied to the assigned credit.
Large language models (LLMs) have long relied on outcome-based reward modeling—evaluating only final answers—leading to insufficient interpretability and robustness in reasoning processes. Method: This paper systematically introduces the Process Reward Modeling (PRM) paradigm, shifting supervision from outcome-level to step- or trajectory-level reasoning evaluation. We establish a comprehensive methodology encompassing data construction, fine-grained reward modeling, test-time scaling, and RLHF integration. Contribution/Results: We empirically validate PRM across diverse domains—including mathematics, code generation, natural language, multimodal reasoning, and robotic agent tasks. To our knowledge, this is the first work to characterize the design space and core challenges of PRMs across multiple domains, releasing a cross-task benchmark, practical implementation guidelines, and open-source resources. Our framework provides foundational theoretical insights, actionable technical pathways, and empirical support for trustworthy reasoning alignment in LLMs.
This study addresses the challenges of sparse rewards and the absence of intermediate process credit assignment in reinforcement learning for deep research agents. We propose a rubric-based fine-grained credit assignment method that leverages task requirements as a shared reference, evaluating the incremental contribution of tool calls by contrasting historical evidence. This approach eliminates reliance on ground-truth answers, enabling process supervision for open-ended tasks. Furthermore, we construct a hybrid reinforcement learning framework that integrates the outcome advantages of GRPO with rubric-based process advantages to guide research decision-making. Experimental results demonstrate that our method comprehensively outperforms open-source baselines across multiple benchmarks. Notably, an 8B-parameter model achieves performance comparable to frontier proprietary models while significantly improving evidence acquisition efficiency.
为解决语言模型代理在长周期任务中评估不足的问题,提出了一种基于人类规则和大语言模型的细化评分方法(GCPC),以更准确地评估代理技能轨迹。
This work addresses the credit assignment challenge in process reward modeling when training relies solely on the correctness of final answers, where individual reasoning steps receive no explicit supervision. The authors propose a Learnable Credit Assignment (LCA) framework that, for the first time, incorporates the “weakest link” principle into outcome-supervised process reward modeling. They formalize the problem as multiple instance learning and introduce a Softmax-weighted sum pooling mechanism to effectively handle strong dependencies and redundancy among reasoning steps. By jointly optimizing credit assignment and reward modeling, LCA significantly outperforms existing outcome-supervised methods across diverse tasks and large language model backbones, demonstrating enhanced capability in identifying reasoning errors and improving overall performance.
This work investigates the capability of large language models (LLMs) in automating fine-grained scoring of mathematical competition proofs—beyond binary correctness assessment—by detecting errors at the step level, classifying their severity, and assigning partial credit. We propose an intelligent agent–based workflow that dynamically generates problem-specific rubrics by integrating reference solution analysis, error localization, and hierarchical penalty rules, enabling multi-step, interpretable scoring. Evaluated on 90 expert-annotated proofs and the MathArena benchmark, our method achieves significantly higher agreement with human graders (Krippendorff’s α increased by 18.3%) and notably improves calibration of partial credit assignment. All code, datasets, and experimental logs are publicly released.
This study investigates how goal-aligned and goal-agnostic reward mechanisms influence decision-making behavior and design diversity in creative tasks. Using a 3D parametric chair design task as the experimental setting, the design process is formalized as a Markov decision process, and a mixed-methods approach—integrating user behavior tracking, experimental psychology paradigms, and hybrid analytical techniques—is employed to systematically examine participants’ exploration strategies and subjective experiences under different reward conditions. The findings reveal that goal-aligned rewards not only enhance goal attainment but also foster more thorough exploration of the design space while preserving diversity. Moreover, the nature of the design goal significantly moderates users’ perceived usefulness of the rewards. These results elucidate the synergistic mechanism between rewards and goals and offer actionable guidelines for designing effective feedback systems in creative design contexts.
本文提出CREDO框架,通过演化语义评分标准与选择性执行信用修正,解决长期语言代理稀疏反馈及中间评估可能错误的问题。
研究提出一种统一方法,通过调整奖励分配来最大化全支付竞赛中的预期总努力,适用于基于排名和基于表现的评分制度。
本文提出ProCredit方法,通过在每一步验证进展并据此分配奖励,解决了长周期任务中基于最终结果奖励导致的训练信号不足问题。
论文提出DARS框架,通过双层信用分配和结构化推理解决基于指令的图像编辑中规划器与渲染器优化效率低的问题。
为了解决多特征自动评分中反馈与评分一致性差的问题,提出了一种统一框架HiFTS,通过生成层次化反馈并结合评分预测来提高评分准确性及反馈质量。