Optimal Multi-Reward Reinforcement Learning

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses multi-reward reinforcement learning in finite-horizon Markov decision processes with unknown transitions, aiming to simultaneously output ε-optimal policies for multiple known reward functions through online interactions alone. To this end, it proposes a novel framework that integrates minimum-variance planning adaptation, conservative evaluation via fresh replay samples, and gap-based multiplicative weight updates. By combining optimistic value estimation with adaptive reward sampling distribution adjustment, the approach enables efficient exploration without additional warm-up costs. The primary contribution lies in establishing a minimax sample complexity upper bound that matches the information-theoretic lower bound, thereby significantly enhancing the efficiency of multi-objective decision-making.
📝 Abstract
We study an unknown-transition finite-horizon Markov decision process (MDP) with a finite collection of known reward functions $\{r^1, r^2, \ldots, r^M\}$. The goal is to output an $ε$-optimal policy for every reward using online episodic interaction only. Performance is measured by the policy error $V_{0}^{*, m} - V_{0}^{\widehatπ^{m}, m}$ where $m\in [M]$ represents the reward function and $V_{0}^{*, m}=\mathbb{E}_{s_1\sim μ}[V_{1}^{*, m}(s_1)]$. Under this setting, we design a provably efficient algorithm to establish a minimax sample complexity bound of $$ O\left(\frac{SAH^3}{ε^2}\log M \mathrm{polylog}\left(\frac{SAH\log M}{\min\left\{ε, 1\right\}δ}\right)\right)$$ episodes, with no additional burn-in cost. This matches the information-theoretic lower bound up to a factor of $ \mathrm{polylog}(SAH\log M/(\min\left\{ε, 1\right\}δ))$. Our method combines three technical ingredients. First, we adapt MVP to reward-switching learning to construct optimistic value estimates. Second, we use fresh replay samples to conservatively evaluate the candidate policies. Third, gap-based multiplicative weights updates adjust the reward-sampling distribution using the differences between these estimates, converting weighted learning progress into simultaneous guarantees for all rewards.
Problem

Research questions and friction points this paper is trying to address.

Multi-Reward Reinforcement Learning
Markov Decision Process
Sample Complexity
Online Episodic Interaction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Reward Reinforcement Learning
Minimax Sample Complexity
Optimistic Value Estimation
Multiplicative Weights Update
Markov Decision Process
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zijun Chen
Department of Computer Science and Engineering, Hong Kong University of Science and Technology
Zihan Zhang
Zihan Zhang
Southern University of Science and Technology
HCI