Score
Designs and implements models that map observations, states, trajectories, or other signals to scalar reward or utility estimates used to train, evaluate, or guide decision-making agents. Builds data collection and preprocessing pipelines for feedback (preferences, ratings, demonstrations), selects representations and learning procedures for reward predictors, and analyzes their calibration, robustness, and susceptibility to mis-specification or reward hacking.
This paper addresses the misalignment between reward models and true objectives in deep reinforcement learning, as well as the resulting limitations in policy optimization. To this end, it introduces— for the first time—a unified taxonomy that systematically organizes reward modeling across three orthogonal dimensions: modeling source (explicit vs. implicit), mechanism design (supervised vs. interactive), and learning paradigm (static vs. dynamic). The survey comprehensively covers mainstream approaches—including inverse reinforcement learning, preference learning, language-model-based feedback, human demonstration distillation, contrastive learning, and online interactive modeling—and critically analyzes evaluation methodologies and practical deployment challenges. This work fills a critical gap in the literature by providing the first systematic, cross-cutting review of reward modeling. It clarifies the technical evolution of the field and identifies four key research frontiers: scalability, generalization, robustness, and human-AI alignment.
Current research on reward modeling and evaluation metrics operates in silos, leading to terminological redundancy, spurious correlations, heightened reward hacking risks, and duplicated efforts in data quality optimization and meta-evaluation. Through a systematic literature review and comparative analysis, we reveal that both reward models and evaluation metrics fundamentally serve the same purpose in language model post-training: preference modeling and performance calibration. Building on this insight, we propose a unified research framework integrating three core directions—preference acquisition, spurious correlation mitigation, and meta-evaluation calibration. Empirical experiments demonstrate that certain evaluation metrics significantly outperform existing reward models on specific tasks. This work clarifies the root causes of conceptual ambiguity in the field and fosters cross-paradigm collaboration, providing both theoretical foundations and practical pathways for developing robust, interpretable, and reusable alignment evaluation systems.
In reinforcement learning, reward function design faces critical challenges including delayed signals, ambiguity, misalignment with task objectives, and induction of undesirable behaviors. To address these, this paper proposes three novel reward mechanisms: teacher-driven, adaptive explainable, and agent-autonomous reward generation. Our core contributions are the first-ever adaptive explainable reward design method and a meta-learning–driven autonomous reward generation framework—enabling a paradigm shift from expert-guided reward specification to online inverse reward modeling by the agent. Technically, we integrate reward shaping, eXplainable AI (XAI)-informed reward modeling, policy-value alignment, and online inverse reward design. Experiments across multiple sparse-reward benchmarks demonstrate over 40% faster training convergence, significantly improved policy robustness, and high reward interpretability—validated by domain experts with 92% inter-rater agreement.
Current post-training of language models relies on abstract scalar rewards, which lack transparency regarding the instructional content of preference data and can lead models to learn spurious correlations, resulting in undesirable behaviors such as excessive stylization or sycophancy. This work proposes a data-centric post-training framework that, for the first time, leverages interpretability methods to explicitly model latent conceptual signals within preference data. By analyzing and identifying key features that distinguish preferred from non-preferred responses prior to optimization, the approach integrates interpretability protocols, statistical hypothesis testing, and fine-grained interventions at both feature and data levels. This enables effective diagnosis and suppression of harmful learning signals, significantly reducing off-target behaviors across multiple benchmarks while enhancing model safety and controllable personality.
Addressing practical challenges in reinforcement learning—such as sparse and delayed rewards and training instability—this paper presents a systematic survey of reward engineering and reward shaping. We propose the first fine-grained taxonomy of reward design techniques, explicitly exposing their implicit assumptions and failure boundaries. Furthermore, we introduce an evaluation framework for reward shaping that jointly balances interpretability and empirical effectiveness. Our analysis integrates theoretical foundations of RL, deep RL practice, formal modeling of reward functions, and cross-domain applications—including robotics and autonomous driving. This work fills a critical gap by providing the first comprehensive, methodology-driven survey of reward design. It establishes a unified tripartite research framework comprising methodology, taxonomic classification, and application boundaries. The resulting synthesis delivers a reproducible, transferable engineering guide for algorithm designers, significantly enhancing the robustness and real-world deployability of RL systems. (149 words)
Existing process reward models (PRMs) rely on costly step-level human annotations or ground-truth reference solutions, limiting their applicability to domains like mathematical reasoning where gold-standard process annotations are unavailable. Method: We propose SPARK, the first framework for ground-truth-free process-level reward modeling. It employs a generator-verifier collaborative paradigm to produce diverse solution paths, integrates parallel self-consistency scoring, sequence-level meta-critique, and chain-of-thought verification (PRM-CoT) to construct synthetic verification data for fine-tuning a generative PRM, and incorporates format constraints to mitigate reward hacking. Contribution/Results: On ProcessBench, SPARK achieves 67.5 F1—surpassing the ground-truth-supervised baseline (66.4). Across six mathematical reasoning benchmarks, it attains a mean accuracy of 47.4%, significantly outperforming RLVR (43.9%) and establishing the first effective process-supervised reinforcement learning method without reference answers.
This work addresses the excessive sensitivity of existing neural reward models to semantically equivalent responses, a flaw that often triggers reward gaming in reinforcement learning and degrades policy performance. To mitigate this issue, the authors propose a training-free discretization method that leverages Monte Carlo Dropout to generate reward clusters, mapping continuous rewards to discrete values while preserving discriminative capacity. To better evaluate reward models, they introduce two novel metrics: discriminability and specificity. Empirical results demonstrate that the proposed approach significantly suppresses reward gaming and enhances policy quality across both control and natural language reinforcement learning environments.
Sparse, delayed, and weakly informative reward signals severely hinder the efficiency of reinforcement learning, and existing reward shaping methods often fail to adapt to dynamic environments. This work proposes the first unified analytical framework encompassing temporal, informational, and theoretical dimensions to systematically categorize and compare twelve classes of dynamic reward shaping and related adaptive mechanisms. It clearly distinguishes between parameter corrections and state-dependent modifications, and precisely delineates the boundaries among additive shaping, reward replacement, and correlated guidance. Through integrated theoretical analysis and taxonomic synthesis, the study identifies conditions under which optimality is preserved in the presence of modern RL components such as experience replay, bootstrapped critics, and reward normalization, and for the first time elucidates the intrinsic relationship between adaptation rate and learning stability.
Current robot learning approaches predominantly rely on sparse success signals at task completion, lacking effective feedback on behavioral progress and thus struggling to acquire complex skills. This work proposes a unified three-dimensional framework encompassing input-output interfaces, modeling mechanisms, and evaluation protocols to systematically categorize existing progress-based reward modeling methods. By clarifying key distinctions in observation spaces, goal representations, supervision sources, and reward generation strategies, the study establishes a coherent taxonomy of the field. Furthermore, through the integration of relevant datasets and standardized evaluation protocols, it constructs comparable benchmarks and identifies fundamental limitations of current approaches, thereby offering a clear roadmap for future research directions in progress-driven robotic learning.
This work addresses the challenge of reward hacking in large language model agents, which can arise from entanglement between internal states and environmental context, rendering risk prediction based solely on internal activations unreliable. To mitigate this, the authors propose a context-calibrated mechanistic monitoring framework that treats reward-hacking activations as latent policy states and integrates token-level entropy with decision-context features to enhance risk prediction accuracy. The approach leverages activation scoring, entropy analysis, context-aware feature extraction, adapter fine-tuning, and activation steering to effectively identify high-risk behaviors. Evaluated on Gameable ALFWorld and WebShop environments, the method significantly outperforms baseline approaches relying exclusively on activation signals, demonstrating improved detection of exploitative strategies and reduced agent exploitation of reward loopholes.