Score
Designs and implements evaluation pipelines, metrics, frameworks and systems to assess models or policies using historical (offline) datasets and live (online) experiments; this includes building offline policy-evaluation estimators, specifying offline and online evaluation metrics, constructing automated offline/online testing and validation pipelines, and analyzing discrepancies between offline and online results to guide model validation and deployment.
Offline evaluation metrics for autonomous driving exhibit low correlation with online performance and fail to anticipate cumulative errors in closed-loop operation. To address this, we propose a novel offline metric grounded in cognitive uncertainty modeling, integrating perception and planning within a unified framework. Our approach leverages large-scale closed-loop simulation and real-world testing on ground-truth-annotated datasets to systematically analyze evaluation biases across multiple existing metrics. Experimental results demonstrate that the proposed metric significantly improves detection of latent high-risk scenarios. It achieves over 13% higher correlation with online safety outcomes than conventional metrics in both simulation and real-world settings—with particularly pronounced gains in real-road testing—thereby validating its strong generalizability and engineering practicality.
Online A/B testing is costly, risky, and time-consuming; offline policy evaluation (OPE) mitigates these issues but suffers from reliability degradation due to poor implementation quality—yet no prior work explores LLM-based agents for automated OPE code optimization. This paper introduces GrowthHacker, the first benchmark for evaluating coding LLM agents in OPE, and proposes a two_agent architecture to systematically study agent-driven OPE code generation, iterative refinement, and closed-loop evaluation. Leveraging real-world datasets built upon Open Bandit Pipeline and Scope-RL, and integrating frameworks including CrewAI and AutoGen, we conduct comprehensive comparative experiments. Results show that two_agent achieves a 106.7% average improvement on positive metrics, 100% reliability, and 45% success rate—substantially outperforming all baselines. This validates LLM agents as effective and feasible “growth hackers” for automating and enhancing OPE robustness and efficiency.
Offline evaluation metrics in recommender systems often exhibit poor correlation with online performance, limiting their reliability for predicting real-world effectiveness. To address this misalignment, we propose a generic Pareto-frontier approximation strategy that jointly calibrates multiple offline metrics (e.g., Recall, NDCG) against multidimensional online metrics (e.g., CTR, CVR, GMV) within a single-model framework—without architectural modifications. Our method is model-agnostic, supports parallel A/B testing, and scales efficiently to industrial settings. Evaluated on large-scale production traffic at OTTO’s e-commerce platform, the approach significantly improves the consistency between offline metric trends and observed online outcomes. It provides an interpretable, reusable, and scientifically grounded foundation for metric selection and algorithmic iteration in industrial recommender systems.
This work investigates how online versus offline data collection strategies affect the generalization and task performance of world models in model-based reinforcement learning. We identify that offline training suffers from insufficient state coverage, leading to out-of-distribution states at test time and substantial performance degradation. To mitigate this, we propose two mechanisms: (1) incorporating limited online interaction—under fixed or adaptive scheduling—to recalibrate the world model; and (2) augmenting offline datasets with exploratory trajectories to improve state coverage. Systematic evaluation across 31 continuous control benchmarks reveals that purely offline agents consistently underperform online baselines; however, even minimal online interaction restores—and often exceeds—their performance. Moreover, injecting exploration data significantly enhances the robustness and generalization of offline agents. Our study establishes a scalable co-design paradigm for data acquisition and model training, effectively balancing data efficiency with world model generalization.
In model-based systems engineering, low experimental data reuse efficiency and excessive redundant experiments hinder digital engineering agility. To address this, this paper proposes a case-based reasoning (CBR)-driven experimental management framework that explicitly integrates domain knowledge. The framework features structured experimental metadata modeling, digital twin–enabled scenario semantic alignment, and an interpretable similarity assessment mechanism to intelligently determine whether historical experiments can be transferred to address new verification queries. Its key innovation lies in embedding domain knowledge explicitly into both the CBR retrieval and adaptation stages, thereby enabling trustworthy cross-operating-condition and cross-configuration experimental data reuse. Evaluated on an industrial-scale vehicle energy system design case, the framework reduces redundant experiments by 37% and shortens early verification cycles by 42% on average, significantly enhancing iterative efficiency in digital engineering and advancing intelligent experimental management.
This work addresses a key challenge in offline-to-online reinforcement learning: how to efficiently select and fine-tune policies under a limited online interaction budget while avoiding performance degradation due to algorithmic or hyperparameter sensitivity. The paper introduces the first active policy selection framework tailored for such constrained settings. By constructing an upper confidence bound based on a local linear performance prediction model, the method dynamically balances resource allocation between online evaluation and fine-tuning, adaptively identifying the most promising policies for optimization. This approach overcomes the limitations of conventional strategies that either deploy a single policy or uniformly distribute the budget across candidates. Empirical results across multiple environments demonstrate that the proposed method significantly outperforms existing baselines, achieving more efficient utilization of scarce online interaction resources.