Tail-Influence Sampling for CVaR Policy Evaluation

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high evaluation cost of Conditional Value-at-Risk (CVaR) estimation in rare failure scenarios. To this end, it proposes a tail-influence sampling method that quantifies component contributions to tail risk for dynamic allocation of a fixed budget. By integrating Bellman reuse, Neyman allocation, and adaptive importance sampling, the approach derives an aggregated metric achieving near-optimal variance efficiency, while introducing a visit-anchored variant to prevent pilot underestimation and ensure robustness. Experimental results demonstrate that the proposed method reduces mean squared error by 41%–76% in the CliffWalking environment and decreases estimation errors to approximately one-third of those produced by mean-oriented baselines in large language model red-teaming tasks.
📝 Abstract
Policies with similar mean returns can differ sharply in rare failures, yet estimating lower-tail conditional value-at-risk (CVaR) accurately can require many costly rollouts. When different conditional components of a stochastic workflow can be queried separately, we ask how to allocate a fixed evaluation budget to estimate a fixed policy's CVaR most accurately. We derive a tail influence for each queryable conditional law that aggregates how its uncertainty affects CVaR across every Bellman reuse. Its variance yields the fixed-design efficiency bound and the oracle Neyman allocation. Tail-Influence Sampling (TIS) estimates these influence scales from a pilot model and reallocates fresh queries toward kernels that matter most for the tail; a visitation-anchored variant protects against pilot underallocation. Under fixed dimension and a positive quantile margin, TIS attains oracle asymptotic variance and first-order MSE including pilot cost, while the anchored variant is within a factor two of the oracle. We also characterize an exact-grid regime in which tail- and mean-optimal allocations coincide. On CliffWalking, TIS reduces MSE by 41% versus learned occupancy and 76% versus complete rollouts at the same charged transition budget. In frozen language-model review workflows, anchored TIS beats an equally regularized mean-influence blend in 23 of 24 MMLU-Pro settings and reaches 2.4-3.4$\times$ lower MSE than rollouts on six-call FinQA reviews.
Problem

Research questions and friction points this paper is trying to address.

CVaR
policy evaluation
budget allocation
tail risk
rare failures
Innovation

Methods, ideas, or system contributions that make the work stand out.

Tail-Influence Sampling
CVaR Policy Evaluation
Neyman Allocation
Bellman Reuse
Budget Allocation