🤖 AI Summary
This study addresses the inefficiency of traditional offline evaluation methods for large language models, where policy and reward distribution shifts necessitate repeated model fitting. To overcome this, we propose PFN-OPE, a framework that reformulates off-policy evaluation (OPE) as an amortizable sequence prediction task for the first time. By adopting a Prior-data Fitted Network (PFN) architecture pretrained on a large-scale pool of contextual bandit tasks, our method enables cross-task amortized evaluation. It yields value estimates via a single forward pass, eliminating the need to retrain models when encountering new logging policies or reward definitions. Extensive experiments on benchmarks such as HelpSteer2 demonstrate that PFN-OPE reduces evaluation error by 2.0× to 9.3× compared to the strongest baselines, substantially improving both computational efficiency and estimation accuracy.
📝 Abstract
Accurate evaluation is central to selecting which LLM to deploy, yet testing a candidate on live traffic exposes real users to an unvetted model. Teams therefore evaluate candidates offline, on data produced by already-deployed models. This is off-policy evaluation (OPE), and it faces two distribution shifts: as a model is updated in post-training, its responses diverge from the logged ones (policy shift), and the reward definition under which it is judged changes with business requirements (reward shift). Classical OPE methods are ill-suited to this continual-deployment setting because they are defined per task and require fitting from scratch on every new logged dataset or reward definition. To address this, we propose PFN-OPE, a prior-data fitted network that amortizes OPE across a distribution of contextual-bandit tasks. We pretrain it once on tasks constructed from a pool of LLM responses scored by several reward functions, in which both shifts occur. At test-time it maps a logged dataset and one sampled target response per prompt to a value estimate in a single forward pass, with no per-task fitting. On HelpSteer2 and UltraFeedback with Qwen, Llama, and Gemma policies, PFN-OPE achieves 2.0 to 9.3 times lower error than the best baselines across all tested configurations in the reward-shifted settings.