🤖 AI Summary
This work addresses a critical limitation in existing reinforcement learning methods based on multinomial logit function approximation, which provide only worst-case regret bounds and overlook the crucial role of variance arising from environmental interactions. To bridge this gap, we propose a computationally efficient algorithm that establishes, for the first time, an explicit variance-adaptive regret bound, enabling a fine-grained characterization of instance-dependent performance. By integrating multinomial logit approximation, variance-aware analysis, and efficient policy optimization, our method achieves instance-optimal rates and substantially narrows the gap between theoretical upper and lower bounds. Empirical evaluations demonstrate that the proposed algorithm significantly outperforms existing approaches in learning optimal policies.
📝 Abstract
Reinforcement learning with multinomial logistic (MNL) function approximation has become an important framework due to its flexibility and broad applicability. While existing studies have established regret guarantees under worst-case analysis, they do not capture how performance depends on the variability of the interaction between the learner and the environment. In this paper, we develop a new theoretical analysis for MNL-based Markov decision processes that yields explicit variance-adaptive regret bounds. Our algorithm is computationally efficient and achieves the instance-wise optimal rate of regret, narrowing the gap between upper and lower bounds. Our numerical experiments validate that our method learns optimal policies more efficiently than conventional approaches.