Amortized Off-Policy Evaluation for LLMs
This study addresses the inefficiency of traditional offline evaluation methods for large language models, where policy and reward distribution shifts necessitate repeated model fitting. To overcome this, we propose PFN-OPE, a framework that reformulates off-policy evaluation (OPE) as an amortizable sequence prediction task for the first time. By adopting a Prior-data Fitted Network (PFN) architecture pretrained on a large-scale pool of contextual bandit tasks, our method enables cross-task amortized evaluation. It yields value estimates via a single forward pass, eliminating the need to retrain models when encountering new logging policies or reward definitions. Extensive experiments on benchmarks such as HelpSteer2 demonstrate that PFN-OPE reduces evaluation error by 2.0× to 9.3× compared to the strongest baselines, substantially improving both computational efficiency and estimation accuracy.