🤖 AI Summary
This study addresses the unclear finite-sample benefits of reward and policy regularization in adversarial imitation learning by proposing a dual-regularization algorithm that combines KL policy regularization with quadratic reward penalties. Methodologically, it employs online mirror descent to handle general convex reward classes, integrating optimistic KL-regularized policy learning with function approximation techniques. Theoretically, this work is the first to simultaneously achieve an O(1/ε) sample complexity for both expert demonstrations and online interactions. It establishes an Õ(1/K + 1/N) bound on the imitation gap under fixed regularization parameters and rigorously characterizes the complementary statistical advantages of the two regularizers. These results provide a solid theoretical foundation for empirically successful methods in this domain.
📝 Abstract
We study adversarial imitation learning (AIL), in which an agent learns to imitate expert demonstrations by optimizing a policy against an adversarial reward that distinguishes expert and learner behavior. Historically, reward regularization and entropy-based policy regularization are key components of empirically successful methods such as GAIL and LS-IQ, yet their finite-sample benefits remain underexplored. We establish fast rates for jointly regularized AIL in finite-horizon Markov decision processes with general function approximation. Our model-free algorithm, Dually Regularized AIL, combines KL policy regularization with a quadratic reward penalty weighted by expert and learner occupancies. With K online episodes and N expert trajectories, we prove a $\widetilde{O}\left(\frac{1}{K}+\frac{1}{N}\right)$ bound on the regularized imitation gap for fixed regularization parameters. Our analysis combines an online mirror descent construction for general convex reward classes to control estimation error from finite expert data and stochastic learner feedback, with a sharp analysis of optimistic KL-regularized policy learning. To the best of our knowledge, Dually Regularized AIL is the first algorithm to simultaneously achieve $\widetilde{O}\left(\frac{1}{\epsilon}\right)$ sample complexity in both expert demonstrations and online interactions for this regularized AIL objective, even with stochastic experts. These results provide a rigorous characterization of the complementary statistical benefits of reward and policy regularization in AIL.