🤖 AI Summary
This work addresses the growing pessimistic bias in off-policy reinforcement learning caused by multi-step returns that rely on suboptimal historical actions. To mitigate this issue, the authors propose Expected Quantile n-step Q-learning (ENQ), which introduces an asymmetric expected quantile loss into multi-step Q-learning for the first time, replacing the conventional symmetric temporal difference loss. Theoretical analysis shows that ENQ eliminates bias within the support of the return distribution under deterministic dynamics and connects to established theoretical bounds for long-horizon Q-learning; under stochastic dynamics, it provides horizon-independent two-sided bias bounds. Empirically, ENQ matches the performance of LQL across 27 manipulation and navigation tasks using a single quantile level (τ=0.8), achieves higher training throughput, and demonstrates substantial gains when combined with an ensemble of ten critics.
📝 Abstract
Multi-step returns accelerate reward propagation in off-policy reinforcement learning, but couple the evaluation of each decision to the suboptimal logged actions that follow it, inducing a pessimistic bias that grows with the horizon. We propose Expectile $n$-step Q-learning (ENQ), which replaces the symmetric $n$-step temporal-difference (TD) loss with an asymmetric expectile loss on the action-value error, with expectile level $τ$ as the only method-specific hyperparameter added beyond $n$-step TD. We prove that the ENQ operator is a $γ^{n}$-contraction. Under deterministic dynamics, at $τ=1$, its bias vanishes at the optimal action-value function $Q^*$ on covered in-support pairs, and the corresponding fixed point satisfies the separation-$n$ instance and its multiples of the lower-bound inequality used by Long-Horizon Q-learning (LQL). Under stochastic dynamics, the operator bias admits two-sided bounds with horizon-independent noise constants. Using a single expectile level $τ=0.8$ and a fixed backup horizon across 27 manipulation and navigation task instances, ENQ is competitive with LQL on aggregate, achieves higher measured training-step throughput in our profiling study, and benefits more from a ten-critic ensemble in a controlled scaling experiment.