model-based evaluation

Designs and implements evaluation procedures that use learned predictive models or simulators to estimate a system’s performance, outcomes, or long-term returns without full redeployment. This includes building predictive user/feedback simulators, adapting doubly robust off‑policy estimators, running simulation‑based test‑time reranking, and applying sequential model‑based optimization to predict and compare candidate outcomes.

model-basedevaluation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.88
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$201K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the cold-start challenge in predictive modeling when deploying new decision policies, where historical data fails to reflect system responses. To overcome this, we propose a Sim2Real transfer framework based on counterfactual simulation. A simulator generates counterfactual trajectories to train predictive models for zero-shot transfer, complemented by a lightweight calibration mechanism leveraging early real-world observations to bridge the sim-to-real gap. The method is validated in a real-world inventory control scenario. Experimental results demonstrate that the simulation-trained model reduces Mean Absolute Percentage Error (MAPE) by 1.2–18.7 percentage points compared to historical data baselines, with calibration further decreasing errors by up to 2.5 percentage points. These findings establish an efficient and practical solution for cold-start prediction under novel decision strategies.

cold-start problemcounterfactual simulationforecasting

This paper challenges the unverified implicit assumption in the predict-then-optimize paradigm that “higher prediction accuracy necessarily yields better downstream decisions,” particularly in multiclass classification settings. Method: We propose a controllable, interpretable multiclass prediction simulation framework that explicitly models error types and distributions, enabling systematic analysis of how classification errors affect decision quality in constrained optimization. Contribution/Results: Experiments on job scheduling and other combinatorial optimization tasks reveal a nonlinear relationship between prediction error and decision performance: improving prediction accuracy does not guarantee improved solution quality—and can even degrade decisions when error patterns shift. Our findings question the conventional coupling logic between prediction and optimization, providing theoretical foundations and practical guidance for designing, evaluating, and calibrating classifiers specifically tailored to decision objectives.

Assessing Predict-Then-Optimize performance in machine scheduling problemsEvaluating how prediction error affects optimization solution qualitySimulating multiclass classifier predictions for experimental analysis

Predicting Long Term Sequential Policy Value Using Softer Surrogates

Dec 30, 2024
HN
H. Nam
🏛️ Stanford University

Traditional offline policy evaluation (OPE) struggles to rapidly and accurately estimate the long-term value of newly deployed policies—such as those in healthcare or education—when they introduce novel actions or operate in previously unseen environments. Method: We propose a transfer-based OPE framework leveraging short-horizon data. It integrates counterfactual reasoning, importance sampling, and doubly robust estimation to construct two novel estimators. For the first time, our approach enables testable, low-variance value transfer from short-horizon to full-horizon evaluation under policy distribution shift. Theoretical analysis establishes consistency and asymptotic normality. Results: Experiments on HIV and sepsis treatment simulators demonstrate that our method achieves statistically discriminative full-horizon value estimates using only 10% of the observational trajectory length—substantially outperforming existing OPE baselines.

Long-term EffectsNovel PoliciesPolicy Evaluation

Offline Model-Based Optimization by Learning to Rank

Oct 15, 2024
RT
Rong-Xi Tan
🏛️ Nanjing University | Hohai University | Huawei Technologies Ltd.

Offline model-based optimization (MBO) faces a fundamental challenge: regression models trained on static datasets suffer from out-of-distribution errors, leading to overestimation of suboptimal designs and misguiding the optimization process. This work observes that MBO’s core objective is to **identify promising design rankings**, not to predict absolute performance scores accurately. Accordingly, we propose the first integration of **Learning to Rank (LTR)** into offline MBO. Instead of minimizing mean squared error, our method employs pairwise or listwise ranking losses within an offline reinforcement learning framework to explicitly model relative design preferences. We further derive a theoretical upper bound on the generalization error of ranking loss in this setting. Evaluated across diverse benchmark tasks, our approach consistently outperforms 20 state-of-the-art methods, achieving superior robustness and higher-quality optimal solutions.

Offline MBO aims to maximize black-box function using fixed dataset.Proposed ranking-based model prioritizes designs by relative scores, improving performance.Regression models often overestimate scores, leading to suboptimal designs.

Post Reinforcement Learning Inference

Feb 17, 2023
VS
Vasilis Syrgkanis
🏛️ Stanford University | Hong Kong University of Science and Technology

In reinforcement learning, adaptive interaction data—where the behavior policy is nonstationary—invalidates standard estimators, undermining asymptotic normality for off-policy counterfactual policy evaluation and dynamic treatment effect (DTE) inference. To address this, we propose a weighted Z-estimation framework that constructs time-varying adaptive weights to stabilize heteroskedasticity, achieving, for the first time in the RL off-policy setting, both consistent and asymptotically normal DTE estimation. Our approach integrates dynamic causal inference with asymptotic statistical theory, enabling rigorous hypothesis testing and construction of uniformly valid confidence regions. Simulation studies and real-world RL experiments demonstrate substantial improvements in confidence interval coverage and statistical power. The method provides the first solution for structural parameter inference under adaptive experimentation that simultaneously offers theoretical guarantees—namely consistency, asymptotic normality, and uniform validity—and empirical robustness.

Address nonstationary variance in adaptive reinforcement learning environmentsDevelop weighted Z-estimation for dynamic treatment effect analysisEstimate counterfactual policies post reinforcement learning data collection

Latest Papers

What's happening recently
View more

This study addresses the sim-to-real gap arising from calibration data confusion and drift in pretrained simulators by investigating how to effectively integrate cheap but biased simulations with expensive yet unbiased real-world experiments in sequential decision-making. The authors extend the simulation lemma to decompose policy value error, revealing that under passive learning, the reachability gap is irreducible. To overcome this limitation, they propose Fisher-SEP, an active experimentation strategy based on the Fisher information matrix, which minimizes the predictive variance of the target policy’s value through Bayesian posterior inference. Empirical validation on two real-world domains—vending machine supply chains and mobile HIV testing—demonstrates that early real-world trials yield substantial long-term benefits in the former, while only active exploration effectively covers low-monitoring regions in the latter.

experimental designpolicy evaluationsequential decision making

This work addresses the high computational cost of high-fidelity simulation models, which hinders efficient training and retraining of reinforcement learning agents in dynamic environments. To overcome this limitation, the paper proposes a learnable surrogate modeling framework tailored for dynamic settings, which approximates the input–output mapping of high-fidelity simulations to substantially reduce training overhead when system dynamics, parameters, or reward structures change. By integrating discrete-event simulation, reinforcement learning, and data-driven surrogate modeling techniques, the framework enables rapid adaptation of policies to environmental shifts. Empirical evaluation in stochastic service systems demonstrates significant acceleration in both initial training and retraining processes, thereby enhancing the adaptability of reinforcement learning policies to evolving conditions.

Discrete-Event SimulationReinforcement LearningSimulation Surrogate Models

This work addresses the challenge of accurately predicting pretraining loss for large language models across varying model scales, batch sizes, and training steps—particularly under dynamically changing batch sizes and extreme extrapolation of compute budgets. To this end, the authors propose a loss prediction model grounded in a noisy quadratic system, which explicitly models test loss as a function of model size $N$, batch size $B$, and number of weight updates $K$. The framework enables joint optimization of training configurations under composite constraints on time, memory, and computational resources. Notably, it achieves high-precision loss prediction in variable-batch settings—a first in the field—and significantly outperforms existing heuristics such as Chinchilla, even when extrapolating up to 1000× beyond observed compute budgets. The recommended $(N, B, K)$ configurations closely align with empirically optimal solutions.

batch sizecompute budgetlarge language models

Fixed-size benchmarking in model evaluation often fails to balance efficiency, statistical reliability, and diverse objectives, leading to either excessive resource consumption or unreliable results. This work proposes the first adaptive framework that integrates sequential testing into AI model evaluation, dynamically allocating evaluation data based on stopping criteria tailored for model ranking and selection tasks. By combining sequential hypothesis testing, minimum detectable effect analysis, and diminishing returns detection, the method achieves substantial gains in efficiency without compromising rigor. Empirical validation on the Open VLM Leaderboard demonstrates an 80% reduction in computational cost while maintaining a confidence interval width of 2.5 points, significantly enhancing both the practicality and scalability of model evaluation.

computational costfixed-size benchmarksmodel evaluation

This work addresses the critical yet underexplored impact of deployment strategies on performance in multi-horizon volatility forecasting. The authors propose an adaptive deployment framework that dynamically selects rollout rules based on validation-set performance, automatically balancing prediction accuracy against inference cost across diverse model architectures—from linear models to PatchTST—and loss functions such as MSE and QLIKE. Experiments on volatility sequences of 20 stocks demonstrate that non-default rollout rules significantly outperform standard multi-output (MIMO) deployment. Notably, a small subset of rollout rules can closely approximate the performance of large ensembles at substantially lower computational overhead, with strategy effectiveness highly dependent on the choice of evaluation metric. This study underscores deployment strategy as a key source of adaptability in financial forecasting, challenging conventional fixed-deployment paradigms.

deployment policymodel adaptivenessmulti-horizon forecasting

Hot Scholars

JZ

James Zou

Stanford University
Machine learningcomputational biologycomputational healthstatistics
JS

Jianwen Sun

Software Engineering Application Technology Lab, Huawei, China
Software engineeringDeep reinforcement learning
CC

Chenhang Cui

National University of Singapore
AI AlignmentFoundation ModelsAI safety
AJ

Adam Jatowt

Professor at Univ. of Innsbruck (previously Kyoto Univ.)
question answeringlarge language modelsinformation retrievalRAG
HY

Huaxiu Yao

Assistant Professor of Computer Science and Data Science, UNC Chapel Hill
Machine LearningFoundation ModelsAI AlignmentAI Agent