Dynamic Budget Allocation for LLM Evaluation under Hard Resource Constraints

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of time-to-event data truncation and budget overruns caused by hard computational resource constraints in evaluating multi-turn interactions with large language models. To this end, we propose HARP, a framework that integrates survival analysis, dynamic programming, and statistical inference. Through adaptive budget reallocation and a reflow mechanism, HARP calibrates lower-bound predictions of time-to-event outcomes and estimates evaluation metrics under strict computational limits. As the first adaptive solution satisfying hard resource constraints, it provides finite-sample coverage guarantees and unbiased estimation properties. Experimental results demonstrate that HARP never exceeds the allocated budget in tasks such as jailbreak detection, while achieving near-nominal coverage rates and low-variance estimates.
📝 Abstract
We evaluate large language models (LLMs) in multi-turn interactions through their time-to-event: the number of interaction steps required to produce an event of interest, such as a successful jailbreak or agentic task completion. Under limited compute, interactions may be terminated before the event occurs, so that event times are only partially observed (censored). Existing allocation methods for calibrating time-to-event bounds satisfy the budget only in expectation and can exceed the available budget on a particular evaluation run. Enforcing a hard constraint is particularly challenging as the cost of a trajectory is initially unknown. We introduce Hard-budget Allocation with Reflow for Predictive calibration (HARP), a budget allocation that satisfies hard resource constraints and adaptively reallocates unused budget. We show how to use HARP to construct lower predictive bounds (LPBs) on the time-to-event and to estimate evaluation metrics such as the jailbreak rate on a fixed benchmark. Although HARP induces dependence in acquisition decisions across different trajectories, we prove that HARP never exceeds the target budget, that its LPBs have finite-sample coverage guarantees, and that its metric estimates are unbiased. Experiments on agentic task success, LLM jailbreaks, toxic content generation, and RAG hallucinations show that HARP achieves coverage close to the nominal level with low variance, while never exceeding the given budget.
Problem

Research questions and friction points this paper is trying to address.

LLM evaluation
dynamic budget allocation
hard resource constraints
time-to-event
censored data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dynamic Budget Allocation
Time-to-Event Evaluation
Lower Predictive Bounds
Hard Resource Constraints
LLM Evaluation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.