Kernel Density Bayesian Inverse Reinforcement Learning

📅 2023-03-13
🏛️ Trans. Mach. Learn. Res.
📈 Citations: 3
✨ Influential: 0
📄 PDF
🤖 AI Summary
In clinical and other few-shot settings, reliably inferring reward functions from extremely limited expert demonstrations remains challenging. This paper proposes a novel framework for Bayesian inverse reinforcement learning (Bayesian IRL) to address this problem. Methodologically, we establish the first posterior concentration guarantee for Bayesian IRL and design a conditional kernel density estimator that relaxes strong assumptions—such as fixed trajectory lengths or structured state spaces—required by conventional approaches. Our contributions are twofold: (1) We theoretically prove that the posterior distribution contracts to the true reward function at the minimax-optimal rate; (2) Empirically, our method achieves significantly faster posterior concentration than existing baselines using only 1–3 demonstrations, demonstrating improved robustness and generalization on both synthetic benchmarks and real-world clinical decision-making tasks.
📝 Abstract
Inverse reinforcement learning (IRL) methods infer an agent's reward function using demonstrations of expert behavior. A Bayesian IRL approach models a distribution over candidate reward functions, capturing a degree of uncertainty in the inferred reward function. This is critical in some applications, such as those involving clinical data. Typically, Bayesian IRL algorithms require large demonstration datasets, which may not be available in practice. In this work, we incorporate existing domain-specific data to achieve better posterior concentration rates. We study a common setting in clinical and biological applications where we have access to expert demonstrations and known reward functions for a set of training tasks. Our aim is to learn the reward function of a new test task given limited expert demonstrations. Existing Bayesian IRL methods impose restrictions on the form of input data, thus limiting the incorporation of training task data. To better leverage information from training tasks, we introduce kernel density Bayesian inverse reinforcement learning (KD-BIRL). Our approach employs a conditional kernel density estimator, which uses the known reward functions of the training tasks to improve the likelihood estimation across a range of reward functions and demonstration samples. Our empirical results highlight KD-BIRL's faster concentration rate in comparison to baselines, particularly in low test task expert demonstration data regimes. Additionally, we are the first to provide theoretical guarantees of posterior concentration for a Bayesian IRL algorithm. Taken together, this work introduces a principled and theoretically grounded framework that enables Bayesian IRL to be applied across a variety of domains.
Problem

Research questions and friction points this paper is trying to address.

Improves reward inference with limited expert demonstrations
Leverages training task data via kernel density estimation
Ensures faster posterior concentration in Bayesian IRL
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses kernel density estimator for likelihood estimation
Incorporates training task data to improve learning
Provides theoretical guarantees for posterior concentration
🔎 Similar Papers
No similar papers found.
Stanford University | University of North Carolina at Chapel Hill | Flatiron Institute | Princeton University | Gladstone Institutes
A
Aishwarya Mandyam
Department of Computer Science, Stanford University
Didong Li
Didong Li
Assistant Professor, Department of Biostatistics, Gillings School of Global Public Health, UNC
Manifold learninggeometric data analysisnonparametric BayesGaussian processesspatial statistics
D
Diana Cai
Flatiron Institute
Andrew Jones
Andrew Jones
Department of Computer Science, Princeton University
B
Barbara E. Engelhardt
Gladstone Institutes, Department of Biomedical Data Science, Stanford University