Dynamic Discrete Choice and Inverse Reinforcement Learning: Inferring Preferences and Beliefs From Human Behavior

📅 2026-08-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文探讨了通过动态离散选择和逆向强化学习方法从人类行为中推断决策者的偏好和信念,解决在动态不确定环境中的决策问题。
📝 Abstract
This article surveys two deeply connected literatures that approach the same fundamental problem from different disciplinary traditions: dynamic discrete choice (DDC) in structural econometrics and inverse reinforcement learning (IRL) in machine learning. Both seek to infer the preferences of decision makers from observed sequential behavior, assuming that individuals act to maximize an expected reward function within a dynamic, uncertain environment formalized as a Markov decision process (MDP). Despite independent origins, the two fields have converged on similar mathematical formulations. We show that the (soft Q-learning) framework now prevalent in IRL is closely related to DDC models under additive extreme value preference shocks, yielding the same softmax (multinomial logit) choice probabilities and smooth Bellman equations that underpin structural estimation in economics. We compare the estimation and computational methods developed in each field. DDC has emphasized maximum likelihood estimation, conditional choice probability estimators, and policy iteration methods. IRL has developed scalable alternatives, including maximum entropy methods, adversarial approaches, and model-free temporal difference estimators that extend to high-dimensional state spaces using deep neural networks. Model-free IRL estimators that combine temporal difference learning with classical two-step methods from econometrics represent a promising direction for bridging the two literatures. Both fields confront shared foundational challenges: the identification problem, whereby multiple reward functions can rationalize the same observed behavior, and the curse of dimensionality in solving the underlying MDP. We believe that cross-fertilization offers substantial opportunities for methodological progress in both fields.
Problem

Research questions and friction points this paper is trying to address.

Dynamic Discrete Choice
Inverse Reinforcement Learning
Markov Decision Process
Preference Inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dynamic Discrete Choice
Inverse Reinforcement Learning
Soft Q-Learning
Temporal Difference Learning
Model-Free Estimators
🔎 Similar Papers
No similar papers found.
P
Pranjal Rawat
Georgetown University
J
John Rust
Professor Emeritus, Georgetown University