offline evaluation

Designs and implements evaluation pipelines, metrics, frameworks and systems to assess models or policies using historical (offline) datasets and live (online) experiments; this includes building offline policy-evaluation estimators, specifying offline and online evaluation metrics, constructing automated offline/online testing and validation pipelines, and analyzing discrepancies between offline and online results to guide model validation and deployment.

offlineevaluation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.75
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$214K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Scalable Offline Metrics for Autonomous Driving

Oct 09, 2025
AA
Animikh Aich
🏛️ Boston University

Offline evaluation metrics for autonomous driving exhibit low correlation with online performance and fail to anticipate cumulative errors in closed-loop operation. To address this, we propose a novel offline metric grounded in cognitive uncertainty modeling, integrating perception and planning within a unified framework. Our approach leverages large-scale closed-loop simulation and real-world testing on ground-truth-annotated datasets to systematically analyze evaluation biases across multiple existing metrics. Experimental results demonstrate that the proposed metric significantly improves detection of latent high-risk scenarios. It achieves over 13% higher correlation with online safety outcomes than conventional metrics in both simulation and real-world settings—with particularly pronounced gains in real-road testing—thereby validating its strong generalizability and engineering practicality.

Bridging offline and online evaluation gap for autonomous drivingDeveloping uncertainty-based offline metrics to predict driving failuresImproving correlation between offline metrics and real-world performance

GrowthHacker: Automated Off-Policy Evaluation Optimization Using Code-Modifying LLM Agents

Nov 02, 2025
JJ
Jie JW Wu
🏛️ Michigan Technological University | Birmingham City University | University of British Columbia

Online A/B testing is costly, risky, and time-consuming; offline policy evaluation (OPE) mitigates these issues but suffers from reliability degradation due to poor implementation quality—yet no prior work explores LLM-based agents for automated OPE code optimization. This paper introduces GrowthHacker, the first benchmark for evaluating coding LLM agents in OPE, and proposes a two_agent architecture to systematically study agent-driven OPE code generation, iterative refinement, and closed-loop evaluation. Leveraging real-world datasets built upon Open Bandit Pipeline and Scope-RL, and integrating frameworks including CrewAI and AutoGen, we conduct comprehensive comparative experiments. Results show that two_agent achieves a 106.7% average improvement on positive metrics, 100% reliability, and 45% success rate—substantially outperforming all baselines. This validates LLM agents as effective and feasible “growth hackers” for automating and enhancing OPE robustness and efficiency.

Automating iterative code optimization cycles for improved OPE performanceOptimizing off-policy evaluation through code-modifying LLM agentsReducing reliance on costly online A/B testing in critical domains

Offline evaluation metrics in recommender systems often exhibit poor correlation with online performance, limiting their reliability for predicting real-world effectiveness. To address this misalignment, we propose a generic Pareto-frontier approximation strategy that jointly calibrates multiple offline metrics (e.g., Recall, NDCG) against multidimensional online metrics (e.g., CTR, CVR, GMV) within a single-model framework—without architectural modifications. Our method is model-agnostic, supports parallel A/B testing, and scales efficiently to industrial settings. Evaluated on large-scale production traffic at OTTO’s e-commerce platform, the approach significantly improves the consistency between offline metric trends and observed online outcomes. It provides an interpretable, reusable, and scientifically grounded foundation for metric selection and algorithmic iteration in industrial recommender systems.

Enable multi-group testing with distinct offline metricsEstablish reliable offline-online metric relationships for recommendersIdentify offline metrics aligning with real-world online impact

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies

Sep 06, 2025
JC
Jiaqi Chen
🏛️ University of Tübingen | ETH Zurich

This work investigates how online versus offline data collection strategies affect the generalization and task performance of world models in model-based reinforcement learning. We identify that offline training suffers from insufficient state coverage, leading to out-of-distribution states at test time and substantial performance degradation. To mitigate this, we propose two mechanisms: (1) incorporating limited online interaction—under fixed or adaptive scheduling—to recalibrate the world model; and (2) augmenting offline datasets with exploratory trajectories to improve state coverage. Systematic evaluation across 31 continuous control benchmarks reveals that purely offline agents consistently underperform online baselines; however, even minimal online interaction restores—and often exceeds—their performance. Moreover, injecting exploration data significantly enhances the robustness and generalization of offline agents. Our study establishes a scalable co-design paradigm for data acquisition and model training, effectively balancing data efficiency with world model generalization.

Addressing state space coverage mismatch between agent imagination and real rolloutsComparing online versus offline data collection strategies in model-based reinforcement learningInvestigating performance degradation caused by Out-Of-Distribution states in offline agents

Reasonable Experiments in Model-Based Systems Engineering

Sep 12, 2025
JC
Johan Cederbladh
🏛️ Mälardalen University | Eindhoven University of Technology | Stellenbosch University | IT University of Copenhagen | University of Oslo | Universidade Federal Rural de Pernambuco | University of Antwerp

In model-based systems engineering, low experimental data reuse efficiency and excessive redundant experiments hinder digital engineering agility. To address this, this paper proposes a case-based reasoning (CBR)-driven experimental management framework that explicitly integrates domain knowledge. The framework features structured experimental metadata modeling, digital twin–enabled scenario semantic alignment, and an interpretable similarity assessment mechanism to intelligently determine whether historical experiments can be transferred to address new verification queries. Its key innovation lies in embedding domain knowledge explicitly into both the CBR retrieval and adaptation stages, thereby enabling trustworthy cross-operating-condition and cross-configuration experimental data reuse. Evaluated on an industrial-scale vehicle energy system design case, the framework reduces redundant experiments by 37% and shortens early verification cycles by 42% on average, significantly enhancing iterative efficiency in digital engineering and advancing intelligent experimental management.

Deciding if existing experiments can answer new engineering questionsIntelligently reusing experiment-related data to avoid redundant experimentsManaging experimental configuration metadata and results efficiently

Latest Papers

What's happening recently
View more

This work addresses a key challenge in offline-to-online reinforcement learning: how to efficiently select and fine-tune policies under a limited online interaction budget while avoiding performance degradation due to algorithmic or hyperparameter sensitivity. The paper introduces the first active policy selection framework tailored for such constrained settings. By constructing an upper confidence bound based on a local linear performance prediction model, the method dynamically balances resource allocation between online evaluation and fine-tuning, adaptively identifying the most promising policies for optimization. This approach overcomes the limitations of conventional strategies that either deploy a single policy or uniformly distribute the budget across candidates. Empirical results across multiple environments demonstrate that the proposed method significantly outperforms existing baselines, achieving more efficient utilization of scarce online interaction resources.

fine-tuninglimited interaction budgetnonstationary domains

Hot Scholars

KG

Kun Gai

Senior Director & Researcher, Alibaba Group
Machine LearningComputational Advertising
PJ

Peng Jiang

Kuaishou Technology
Recommender SystemMachine LearningComputational Advertising
YH

Yao Hu

浙江大学
Machine Learning
JP

Junwei Pan

Tencent, Yahoo Research
Computational AdvertisingRecommendation SystemDeep Learning