Score
Design, implement, and analyze standardized protocols and tooling for evaluating policies from logged data under offline reinforcement learning, including the construction and validation of Fitted Q Evaluation (FQE) estimators and related off‑policy value estimators. This competence covers defining reproducible metrics and baselines, uncertainty and bias diagnostics, cross‑dataset comparison procedures, and reporting conventions that enable comparable, reliable off‑policy policy evaluation.
In offline reinforcement learning, selecting the optimal value function or dynamics model from a candidate set for accurate off-policy evaluation (OPE) of a target policy remains an open challenge lacking theoretical guarantees. Method: We propose the first unified, theoretically grounded OPE selection framework—LSTD-Tournament—featuring dual model-agnostic and model-based selection pathways. It integrates Least-Squares Temporal Difference (LSTD), Fitted Q-Evaluation (FQE), importance sampling, and dynamics fitting within a tournament-style selection mechanism, accompanied by a rigorous bias-variance trade-off analysis. We further introduce a reproducible experimental protocol enabling controlled model misspecification and stable candidate model generation. Results: On Gym benchmarks, LSTD-Tournament reduces average OPE estimation error by 37%, significantly improving selection stability and accuracy. It constitutes the first OPE model selection approach with both provable theoretical guarantees and empirical superiority.
This paper addresses the lack of a unified theoretical framework for distributed off-policy evaluation (OPE). To this end, we propose the Fitted Distributional Evaluation (FDE) framework, which generalizes fitted Q-evaluation to return distribution estimation. Leveraging properties of the distributional Bellman operator, FDE integrates distributional reinforcement learning, offline policy evaluation, and function approximation to yield a provably convergent iterative algorithm applicable to non-tabular settings. FDE is the first systematic framework to formalize design principles for distributional OPE methods, unifying existing disparate approaches and filling a critical gap in principled theoretical foundations for distributional OPE. Experiments on LQR and Atari benchmarks demonstrate that FDE achieves significantly higher estimation accuracy and stability than state-of-the-art methods, validating its effectiveness and generalizability.
Offline reinforcement learning (Offline RL) faces two key challenges: poor policy generalization and distributional shift—existing methods tend to overestimate out-of-distribution actions and require hyperparameter re-tuning for cross-task or cross-dataset transfer. To address this, we propose a dynamic policy switching mechanism at evaluation time, which adaptively fuses an offline RL policy with a behavior cloning policy based on dual uncertainty estimates: epistemic uncertainty (quantified via Monte Carlo Dropout) and aleatoric uncertainty (estimated via data density modeling). Our approach enables zero-shot cross-task transfer without parameter tuning and supports safe, zero-sample fine-tuning from offline to online settings. Evaluated on standard Offline RL benchmarks, it consistently outperforms state-of-the-art methods, achieving faster fine-tuning convergence and superior final performance.
This study addresses the inefficiency of traditional offline evaluation methods for large language models, where policy and reward distribution shifts necessitate repeated model fitting. To overcome this, we propose PFN-OPE, a framework that reformulates off-policy evaluation (OPE) as an amortizable sequence prediction task for the first time. By adopting a Prior-data Fitted Network (PFN) architecture pretrained on a large-scale pool of contextual bandit tasks, our method enables cross-task amortized evaluation. It yields value estimates via a single forward pass, eliminating the need to retrain models when encountering new logging policies or reward definitions. Extensive experiments on benchmarks such as HelpSteer2 demonstrate that PFN-OPE reduces evaluation error by 2.0× to 9.3× compared to the strongest baselines, substantially improving both computational efficiency and estimation accuracy.
This work addresses offline reinforcement learning under unobserved confounding, where input mismatch between behavior and target policies induces biased policy evaluation. We propose the first deep RL algorithm that integrates worst-case environmental robustness into the offline DQN framework. Our method unifies causal counterfactual robust optimization with minimax policy search, yielding a provably safe reward shaping technique that requires neither observability of confounders nor domain-specific priors. Evaluated on 12 confounded Atari benchmarks, our approach consistently outperforms standard DQN—achieving 37%–62% gains in policy performance under severe input mismatch. The framework establishes a new paradigm for safe and reliable offline policy learning in high-dimensional, nonstationary environments.
本文提出了一种基于模型的自助法框架,用于在有限时域、非时齐马尔可夫决策过程中量化离线策略评估的不确定性,提高了数据格式灵活性和统计效率。
This work addresses off-policy evaluation in finite-horizon Markov decision processes under function approximation and limited data coverage. It proposes a novel method based on recursive reweighting and moment matching, which optimizes scalar weights through a value-function discriminator class in a top-down manner to align the reweighted returns with the expected return under the target policy. The approach unifies and generalizes existing techniques such as importance sampling and linear fitted Q-evaluation. Notably, under the sole assumption that the true Q-function is realizable within the chosen function class, it establishes the first finite-sample error bound that is independent of both the statistical complexity of the function class and the ambient dimensionality. This result advances the theoretical understanding of coverage conditions in offline reinforcement learning and significantly enhances the accuracy and robustness of policy evaluation.
Offline reinforcement learning often suffers from value overestimation or excessive pessimism due to distributional shift, and existing conservative methods struggle to surpass the performance of the behavior policy. This work proposes CPQL, a model-free multi-step algorithm that, for the first time, integrates Peng’s Q(λ) operator into conservative Q-learning to replace the standard Bellman operator. This approach implicitly regularizes policy updates while fully leveraging offline trajectories. Theoretical analysis and empirical results demonstrate that CPQL effectively mitigates over-pessimism and overcomes the longstanding trade-off between performance improvement and near-optimality guarantees. On the D4RL benchmark, CPQL significantly outperforms single-step baselines, and its pretrained Q-function enables efficient online fine-tuning—avoiding initial performance degradation and achieving robust gains.
本文提出了一种新的框架,通过稀疏加性结构的非线性函数类来解决强化学习中样本轨迹有限情况下的策略评估问题。
通过分析160,000次训练运行,研究了离线策略学习中的算法性能、超参数敏感性及环境适应性问题,提出了基于数据集的推荐系统以提高研究可靠性。