A Lecture Note on Offline RL and IRL, Part II: Foundations of Inverse Reinforcement Learning and Dynamic Discrete Choice Models

📅 2026-05-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the recovery of implicit reward functions from offline expert demonstrations—formally known as inverse reinforcement learning (IRL)—and establishes, for the first time, a rigorous equivalence between dynamic discrete choice (DDC) models in structural econometrics and entropy-regularized IRL frameworks in machine learning. By unifying classical econometric approaches such as Rust’s nested fixed point algorithm, Hotz–Miller conditional choice probabilities, and Adusumilli–Eckardt temporal-difference methods with modern machine learning techniques including adversarial IRL, occupancy measure matching, and IQ-Learn, the paper constructs a cohesive framework that integrates identification theory with computational practice. This work systematically clarifies the identification conditions, applicability boundaries, and limitations of these diverse methodologies, thereby providing a solid theoretical foundation and practical algorithmic guidance for offline IRL.
📝 Abstract
In the forward reinforcement-learning problem, the reward is fixed and known; the learner is asked to find a good policy or value function. Here we turn the question around. Given offline data generated by an expert, can we recover the reward the expert was optimizing? This is the inverse reinforcement learning problem, and remarkably, two communities, structural econometricians studying dynamic discrete choice (DDC) and machine learners studying entropy-regularized IRL, have been working on exactly the same probabilistic model under different names. We begin by proving their equivalence. We then develop the classical identification result of Magnac and Thesmar and the classical computational paradigms that grew out of it: Rust's nested fixed-point algorithm, the conditional-choice-probability approach of Hotz and Miller, and the two temporal-difference approaches of Adusumilli and Eckardt: linear semi-gradient TD and approximate value iteration. Each route has its limits: dimensionality, transition-kernel estimation, the deadly triad, or projected fixed-point bias. We then walk through the modern ML/IRL strand: adversarial IRL, occupancy matching, IQ-Learn, and offline ML-IRL, deriving each method's actual objective and stating precisely what it does and does not identify. We close with the empirical-risk-minimization framework of Kang et al., which yields a gradient-based estimator for offline IRL/DDC.
Problem

Research questions and friction points this paper is trying to address.

Inverse Reinforcement Learning
Dynamic Discrete Choice
Offline RL
Reward Recovery
Structural Estimation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Inverse Reinforcement Learning
Dynamic Discrete Choice
Offline Learning
Entropy Regularization
Empirical Risk Minimization
💼 Related Jobs
No related jobs found.
E
Enoch Hyunwook Kang
University of Washington, Foster School of Business