Score
Designs and implements algorithms and probabilistic models that infer an agent’s latent reward or objective function from observed sequential decision-making data (demonstrations, trajectories, or action sequences), recovering interpretable reward parameters or latent objectives that explain behavior. Practically this includes formulating inverse-optimal-control or Bayesian inference procedures, modeling dynamics and noise (including nonlinear responses), addressing identifiability and preference shifts, and benchmarking inferred rewards against baselines.
This paper addresses the misalignment between reward models and true objectives in deep reinforcement learning, as well as the resulting limitations in policy optimization. To this end, it introduces— for the first time—a unified taxonomy that systematically organizes reward modeling across three orthogonal dimensions: modeling source (explicit vs. implicit), mechanism design (supervised vs. interactive), and learning paradigm (static vs. dynamic). The survey comprehensively covers mainstream approaches—including inverse reinforcement learning, preference learning, language-model-based feedback, human demonstration distillation, contrastive learning, and online interactive modeling—and critically analyzes evaluation methodologies and practical deployment challenges. This work fills a critical gap in the literature by providing the first systematic, cross-cutting review of reward modeling. It clarifies the technical evolution of the field and identifies four key research frontiers: scalability, generalization, robustness, and human-AI alignment.
This work addresses the limitation of reward inference methods that rely on strong assumptions about human behavior—such as optimality or predefined bias structures—by proposing a robust framework capable of inferring the true reward function from arbitrary human behavior, including suboptimal actions, errors, or goal-communicative demonstrations. Methodologically, it establishes supervised learning as a unified framework for reward inference for the first time, and theoretically proves its asymptotic Bayesian optimality under mild assumptions. The approach integrates Bayesian inference with supervised learning and is validated in robotic manipulation simulations: using only diverse suboptimal demonstrations, it efficiently and accurately recovers reward functions while exhibiting strong generalization. Crucially, the framework eliminates the need to pre-specify or model the underlying behavioral policy, thereby significantly enhancing practical applicability in real-world settings.
This paper addresses three core challenges in inverse reinforcement learning (IRL): (1) reconstructing utility functions from observed behavior; (2) distinguishing whether an agent is a rational, Bayesian utility maximizer—or instead exhibits rational inattention; and (3) online tracking of time-varying utility functions. We propose a unified framework integrating Afriat’s theorem from revealed preference theory with stochastic optimization via an adaptive Langevin dynamics algorithm, augmented by Bayesian inference and inverse stopping-time modeling to ensure robust identification under noise. Our key contributions include the first IRL method capable of identifying cognitive radar and Bayesian sequential detectors, and the novel incorporation of discrete choice models and passive stochastic optimization into time-varying utility tracking. Experiments validate the framework’s effectiveness in detecting constrained utility-maximizing behavior, designing statistical detectors, and reverse-engineering dynamic decision-making systems.
This work addresses the long-standing limitation of inverse reinforcement learning (IRL), which assumes risk-neutral agents and thus fails to infer true risk preferences from behavioral demonstrations. We propose utility learning (UL) as a novel paradigm: for the first time, we explicitly model and infer risk-sensitive utility functions—within the Markov decision process (MDP) framework—directly from expert demonstrations. Theoretically, we establish a partial identifiability theory for utility functions and prove finite-sample convergence of our algorithms. Methodologically, we design two efficient UL algorithms with provable sample complexity guarantees. Empirical results demonstrate that our approach accurately recovers the structural form of the underlying utility function and precisely characterizes human subjects’ risk attitudes—distinguishing between risk aversion and risk seeking—with significant improvements over conventional IRL models.
This work addresses the problem of generalizing expert demonstrations to novel environments or constraint settings. Inverse reinforcement learning (IRL) recovers reward functions to replicate expert behavior, yet suffers from inherent ill-posedness: infinitely many reward functions can rationalize the same observed behavior. To resolve this ambiguity, we propose regularizing IRL by selecting the centroid of the feasible reward set—a principled criterion ensuring policy consistency and interpretability. We derive, for the first time, a closed-form solution for the centroid over a bounded reward subset. Leveraging offline expert demonstrations, we design an efficient algorithm to estimate this centroid and employ it for downstream planning. Experiments demonstrate that our approach significantly improves cross-environment behavioral generalization, faithfully recovering expert intent while providing theoretical convergence guarantees.
In clinical and other few-shot settings, reliably inferring reward functions from extremely limited expert demonstrations remains challenging. This paper proposes a novel framework for Bayesian inverse reinforcement learning (Bayesian IRL) to address this problem. Methodologically, we establish the first posterior concentration guarantee for Bayesian IRL and design a conditional kernel density estimator that relaxes strong assumptions—such as fixed trajectory lengths or structured state spaces—required by conventional approaches. Our contributions are twofold: (1) We theoretically prove that the posterior distribution contracts to the true reward function at the minimax-optimal rate; (2) Empirically, our method achieves significantly faster posterior concentration than existing baselines using only 1–3 demonstrations, demonstrating improved robustness and generalization on both synthetic benchmarks and real-world clinical decision-making tasks.
This work addresses the limitation of conventional inverse reinforcement learning (IRL) methods, which only recover an average reward function from demonstrations generated by multiple experts with potentially heterogeneous reward structures. To overcome this, the paper proposes the first nonparametric Bayesian IRL approach, modeling the reward function using a Dirichlet process prior. By integrating the Chinese Restaurant Process and Gibbs sampling, the method automatically infers both the number and structure of latent reward types without requiring pre-specified cluster counts. To enhance computational efficiency, the authors design a Ray-based data-parallel Gibbs sampling algorithm. Experimental results on the ObjectWorld benchmark demonstrate that the approach accurately recovers two distinct reward types (achieving an Adjusted Rand Index of 1.000) and reliably identifies the correct number of clusters when three reward types are present. An 8-core parallel implementation yields a 4.79× speedup over the sequential version.
This study investigates the probabilistic foundations of reasoning-based decision-making in reinforcement learning, aiming to clarify the relationship between reward signals and Bayesian posterior updates. By introducing KL-regularized soft updates, the work demonstrates that under specific conditions such updates can be rigorously interpreted as Bayesian posteriors within a single fixed probabilistic model, where behavioral changes are entirely driven by observed evidence. Integrating Bayesian inference, information-theoretic channel modeling, and probabilistic graphical models, the approach establishes that posterior updates uniquely determine relative and context-dependent incentive signals, while imposing a consistency constraint on value continuation across update directions. Key contributions include an identifiability theorem for reward representations, a characterization of the mechanistic limits through which posterior updates influence behavior, and the derivation of coherence conditions that reward descriptions must satisfy under multi-directional updates.
Traditional inverse reinforcement learning (IRL) relies on the strong assumption that observed behaviors are nearly optimal, making it ill-suited for scenarios where agents—human or artificial—are still learning and adapting. This work formally introduces and systematically investigates the problem of inferring an agent’s underlying reward function from the behavior of an online learner, modeling the agent either as a no-regret learner or as one asymptotically converging to a Boltzmann-optimal policy. By integrating inverse reinforcement learning with online learning theory and Boltzmann policy modeling, we develop a preference learning framework endowed with theoretical guarantees. We characterize the learnability boundaries of our approach under different learning models, establishing theoretical feasibility in certain settings while revealing fundamental limitations in others.
This study investigates the recovery of implicit reward functions from offline expert demonstrations—formally known as inverse reinforcement learning (IRL)—and establishes, for the first time, a rigorous equivalence between dynamic discrete choice (DDC) models in structural econometrics and entropy-regularized IRL frameworks in machine learning. By unifying classical econometric approaches such as Rust’s nested fixed point algorithm, Hotz–Miller conditional choice probabilities, and Adusumilli–Eckardt temporal-difference methods with modern machine learning techniques including adversarial IRL, occupancy measure matching, and IQ-Learn, the paper constructs a cohesive framework that integrates identification theory with computational practice. This work systematically clarifies the identification conditions, applicability boundaries, and limitations of these diverse methodologies, thereby providing a solid theoretical foundation and practical algorithmic guidance for offline IRL.
本文探讨了通过动态离散选择和逆向强化学习方法从人类行为中推断决策者的偏好和信念,解决在动态不确定环境中的决策问题。