Variance-Averse $n$-Step Offline Reinforcement Learning for Sparse Long-Horizon Environments

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of data heterogeneity in generative offline reinforcement learning, which leads to high-variance policy reproduction and unreliable actions. To this end, it proposes VAN-Flow, a framework integrating a distributional critic with a flow-matching actor. The core innovation lies in a variance-averse operator that smoothly reweights atomic probabilities, enabling unbiased and reliable action selection without hard truncation or auxiliary penalty terms. Additionally, a rejection sampling mechanism is introduced to further enhance generation quality. Evaluated across over forty tasks on the D4RL and OGBench benchmarks, VAN-Flow significantly outperforms existing baselines, demonstrating particularly strong performance in long-horizon and high-variance scenarios.
📝 Abstract
Generative actors are transforming offline reinforcement learning (RL) by enabling expressive policy classes that model complex action distributions. However, this expressiveness also exposes a key challenge in heterogeneous datasets: generative policies can reproduce unreliable action modes whose return distributions exhibit high variance, occasionally yielding high returns by chance but lacking consistency. Consequently, maximizing the expected $Q$-value alone is insufficient for identifying reliable actions. We propose VAN-Flow (Variance-Averse $n$-step Flow), a framework that promotes reliable actions in generative offline RL. VAN-Flow combines (i) a categorical distributional critic, (ii) a variance-averse expectation operator that smoothly reweights atom probabilities to favor actions with both high returns and low dispersion, and (iii) a flow-matching generative actor guided via rejection sampling. Unlike CVaR or mean-variance objectives, the operator redistributes probability mass over the categorical return distribution without hard truncation or auxiliary penalty terms. Across more than 40 tasks from D4RL and OGBench, VAN-Flow consistently outperforms strong baselines, with the largest gains in long-horizon and high-variance regimes where reliable action selection becomes critical.
Problem

Research questions and friction points this paper is trying to address.

offline reinforcement learning
generative policy
variance-averse
heterogeneous datasets
long-horizon
Innovation

Methods, ideas, or system contributions that make the work stand out.

Offline Reinforcement Learning
Variance-Averse Expectation
Categorical Distributional Critic
Flow Matching
Generative Actor
💼 Related Jobs
No related jobs found.
G
Guhyeon Kang
Department of Electrical and Computer Engineering, Sungkyunkwan University, Republic of Korea
Minhae Kwon
Minhae Kwon
Associate Professor at Soongsil University
Reinforcement LearningComputational NeuroscienceAutonomous DrivingFederated Learning