standardise offline rl evaluation

Design, implement, and analyze standardized protocols and tooling for evaluating policies from logged data under offline reinforcement learning, including the construction and validation of Fitted Q Evaluation (FQE) estimators and related off‑policy value estimators. This competence covers defining reproducible metrics and baselines, uncertainty and bias diagnostics, cross‑dataset comparison procedures, and reporting conventions that enable comparable, reliable off‑policy policy evaluation.

standardiseofflinerlevaluation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.11
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Model Selection for Off-policy Evaluation: New Algorithms and Experimental Protocol

Feb 11, 2025
PL
Pai Liu
🏛️ University of Science and Technology of China | University of Illinois Urbana-Champaign | Indraprastha Institute of Information Technology Delhi

In offline reinforcement learning, selecting the optimal value function or dynamics model from a candidate set for accurate off-policy evaluation (OPE) of a target policy remains an open challenge lacking theoretical guarantees. Method: We propose the first unified, theoretically grounded OPE selection framework—LSTD-Tournament—featuring dual model-agnostic and model-based selection pathways. It integrates Least-Squares Temporal Difference (LSTD), Fitted Q-Evaluation (FQE), importance sampling, and dynamics fitting within a tournament-style selection mechanism, accompanied by a rigorous bias-variance trade-off analysis. We further introduce a reproducible experimental protocol enabling controlled model misspecification and stable candidate model generation. Results: On Gym benchmarks, LSTD-Tournament reduces average OPE estimation error by 37%, significantly improving selection stability and accuracy. It constitutes the first OPE model selection approach with both provable theoretical guarantees and empirical superiority.

Development of new experimental evaluation protocolHyperparameter tuning for off-policy evaluationSelection of candidate value functions or dynamics

A Principled Path to Fitted Distributional Evaluation

Jun 24, 2025
SH
Sungee Hong
🏛️ Texas A&M University | University of Texas at Dallas | George Washington University

This paper addresses the lack of a unified theoretical framework for distributed off-policy evaluation (OPE). To this end, we propose the Fitted Distributional Evaluation (FDE) framework, which generalizes fitted Q-evaluation to return distribution estimation. Leveraging properties of the distributional Bellman operator, FDE integrates distributional reinforcement learning, offline policy evaluation, and function approximation to yield a provably convergent iterative algorithm applicable to non-tabular settings. FDE is the first systematic framework to formalize design principles for distributional OPE methods, unifying existing disparate approaches and filling a critical gap in principled theoretical foundations for distributional OPE. Experiments on LQR and Atari benchmarks demonstrate that FDE achieves significantly higher estimation accuracy and stability than state-of-the-art methods, validating its effectiveness and generalizability.

Develops new methods with convergence analysis and theoretical justificationExtends fitted-Q evaluation to distributional off-policy reinforcement learningProvides principles for designing fitted distributional evaluation methods

Evaluation-Time Policy Switching for Offline Reinforcement Learning

Mar 15, 2025
NS
Natinael Solomon Neggatu
🏛️ University of Warwick | Nanyang Technological University

Offline reinforcement learning (Offline RL) faces two key challenges: poor policy generalization and distributional shift—existing methods tend to overestimate out-of-distribution actions and require hyperparameter re-tuning for cross-task or cross-dataset transfer. To address this, we propose a dynamic policy switching mechanism at evaluation time, which adaptively fuses an offline RL policy with a behavior cloning policy based on dual uncertainty estimates: epistemic uncertainty (quantified via Monte Carlo Dropout) and aleatoric uncertainty (estimated via data density modeling). Our approach enables zero-shot cross-task transfer without parameter tuning and supports safe, zero-sample fine-tuning from offline to online settings. Evaluated on standard Offline RL benchmarks, it consistently outperforms state-of-the-art methods, achieving faster fine-tuning convergence and superior final performance.

Adapting offline RL algorithms to diverse tasks without hyper-parameter tuning.Enhancing offline to online fine-tuning using epistemic uncertainty metrics.Overcoming over-estimation of out-of-distribution actions in offline RL.

This study addresses the inefficiency of traditional offline evaluation methods for large language models, where policy and reward distribution shifts necessitate repeated model fitting. To overcome this, we propose PFN-OPE, a framework that reformulates off-policy evaluation (OPE) as an amortizable sequence prediction task for the first time. By adopting a Prior-data Fitted Network (PFN) architecture pretrained on a large-scale pool of contextual bandit tasks, our method enables cross-task amortized evaluation. It yields value estimates via a single forward pass, eliminating the need to retrain models when encountering new logging policies or reward definitions. Extensive experiments on benchmarks such as HelpSteer2 demonstrate that PFN-OPE reduces evaluation error by 2.0× to 9.3× compared to the strongest baselines, substantially improving both computational efficiency and estimation accuracy.

Distribution ShiftLarge Language ModelsOff-Policy Evaluation

Automatic Reward Shaping from Confounded Offline Data

May 16, 2025
ML
Mingxuan Li
🏛️ Columbia University | Syracuse University

This work addresses offline reinforcement learning under unobserved confounding, where input mismatch between behavior and target policies induces biased policy evaluation. We propose the first deep RL algorithm that integrates worst-case environmental robustness into the offline DQN framework. Our method unifies causal counterfactual robust optimization with minimax policy search, yielding a provably safe reward shaping technique that requires neither observability of confounders nor domain-specific priors. Evaluated on 12 confounded Atari benchmarks, our approach consistently outperforms standard DQN—achieving 37%–62% gains in policy performance under severe input mismatch. The framework establishes a new paradigm for safe and reliable offline policy learning in high-dimensional, nonstationary environments.

Improving performance over DQN in confounded Atari game scenariosLearning policies from biased offline data with unobserved confoundersProposing a robust deep RL algorithm for confounded environments

Latest Papers

What's happening recently
View more

This work addresses off-policy evaluation in finite-horizon Markov decision processes under function approximation and limited data coverage. It proposes a novel method based on recursive reweighting and moment matching, which optimizes scalar weights through a value-function discriminator class in a top-down manner to align the reweighted returns with the expected return under the target policy. The approach unifies and generalizes existing techniques such as importance sampling and linear fitted Q-evaluation. Notably, under the sole assumption that the true Q-function is realizable within the chosen function class, it establishes the first finite-sample error bound that is independent of both the statistical complexity of the function class and the ambient dimensionality. This result advances the theoretical understanding of coverage conditions in offline reinforcement learning and significantly enhances the accuracy and robustness of policy evaluation.

finite-horizon MDPsmoment matchingoff-policy evaluation

Offline reinforcement learning often suffers from value overestimation or excessive pessimism due to distributional shift, and existing conservative methods struggle to surpass the performance of the behavior policy. This work proposes CPQL, a model-free multi-step algorithm that, for the first time, integrates Peng’s Q(λ) operator into conservative Q-learning to replace the standard Bellman operator. This approach implicitly regularizes policy updates while fully leveraging offline trajectories. Theoretical analysis and empirical results demonstrate that CPQL effectively mitigates over-pessimism and overcomes the longstanding trade-off between performance improvement and near-optimality guarantees. On the D4RL benchmark, CPQL significantly outperforms single-step baselines, and its pretrained Q-function enables efficient online fine-tuning—avoiding initial performance degradation and achieving robust gains.

behavior policyconservative value estimationmulti-step operator

Hot Scholars

SL

Sergey Levine

UC Berkeley, Physical Intelligence
Machine LearningRoboticsReinforcement Learning
DT

Dacheng Tao

Nanyang Technological University
artificial intelligencemachine learningcomputer visionimage processing
JL

Jiafei Lyu

PhD of Control Science and Engineering, Tsinghua University
deep reinforcement learning
JZ

Junqiao Zhao

Department of Computer science and technology, Tongji University
SLAMLocalizationReinforcement LearningAutonomous Driving