Nonparametric Bayesian Inverse Reinforcement Learning with Data-Parallel Gibbs Sampling

📅 2026-07-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitation of conventional inverse reinforcement learning (IRL) methods, which only recover an average reward function from demonstrations generated by multiple experts with potentially heterogeneous reward structures. To overcome this, the paper proposes the first nonparametric Bayesian IRL approach, modeling the reward function using a Dirichlet process prior. By integrating the Chinese Restaurant Process and Gibbs sampling, the method automatically infers both the number and structure of latent reward types without requiring pre-specified cluster counts. To enhance computational efficiency, the authors design a Ray-based data-parallel Gibbs sampling algorithm. Experimental results on the ObjectWorld benchmark demonstrate that the approach accurately recovers two distinct reward types (achieving an Adjusted Rand Index of 1.000) and reliably identifies the correct number of clusters when three reward types are present. An 8-core parallel implementation yields a 4.79× speedup over the sequential version.
📝 Abstract
Inverse Reinforcement Learning recovers reward functions from expert demonstrations, but standard formulations assume that all demonstrations come from a single expert. When demonstrations are pooled from multiple experts with distinct preferences, parametric methods recover an averaged reward that fits no individual expert well. We implement Nonparametric Bayesian Inverse Reinforcement Learning with a Dirichlet Process prior over reward functions, allowing the number of latent reward types to be inferred jointly with the rewards themselves. Inference uses a collapsed Gibbs sampler combining a Chinese Restaurant Process update for cluster assignments with a Metropolis-Hastings update for reward weights, and soft value iteration as the inner planning routine. We evaluate on a 10x10 ObjectWorld grid with two and three ground-truth reward types. The serial sampler recovers K=2 with Adjusted Rand Index of 1.000, substantially outperforming a Maximum Entropy IRL baseline (ARI=0.000). Extension to K=3 shows that the sampler correctly identifies the number of clusters in all runs; assignment ARI of 0.48-0.58 reflects behavioral overlap between expert types that persists across grid instantiations, revealing that reliable K=3 evaluation on ObjectWorld requires controlled object placement rather than random seeding. We further parallelize the sampler across CPU cores using Ray on HPC hardware, achieving a peak speedup of 4.79x at 8 workers, and characterize a throughput-versus-accuracy tradeoff arising from the consensus merge heuristic used during state aggregation. Code and a containerized environment are available at https://github.com/dasashreeya/np_bayes_irl.
Problem

Research questions and friction points this paper is trying to address.

Inverse Reinforcement Learning
Nonparametric Bayesian
Multiple Experts
Reward Function
Dirichlet Process
Innovation

Methods, ideas, or system contributions that make the work stand out.

Nonparametric Bayesian IRL
Dirichlet Process
Collapsed Gibbs Sampling
Data-Parallel Inference
Reward Clustering
🔎 Similar Papers
No similar papers found.
S
Sai Anirudh Katupilla
University of Maryland, College Park
S
Shreeya Dasa Lakshminath
University of Maryland, College Park