Test-Time Scaling via Budgeted Multi-Attribute Verification

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenging joint decision problem of candidate and attribute selection in LLM answer verification under a shared budget, proposing the BMA-GAI algorithm. This method integrates cost-aware selection, adaptive sampling, and anytime-valid statistical inference to enable efficient certification by eliminating the need for an independent confirmation phase. Theoretically, it establishes asymptotic coverage guarantees and matches the information-theoretic lower bound, thereby proving its first-order optimality. Empirically, evaluations on synthetic benchmarks and LLM tasks demonstrate that BMA-GAI allocates budgets more effectively than existing methods, certifying a greater number of high-quality candidates.
📝 Abstract
Verifying LLM-generated answers under a shared computational budget requires jointly deciding which candidates to inspect and which verification attributes to evaluate. We formulate this problem as multi-attribute good-arm identification under a global budget: each candidate is an arm evaluated along several costly attributes, and the goal is to certify as many candidates as possible whose mean scores exceed the prescribed thresholds on all attributes. We propose \textsc{BMA-GAI}, an algorithm that combines cost-aware arm selection with adaptive sampling of attributes. Every observation serves both to guide adaptive allocation and to support anytime-valid certification, which removes the need for a separate confirmation stage. We establish an asymptotic coverage guarantee for \textsc{BMA-GAI} and derive a matching information-theoretic converse that characterizes the intrinsic complexity of the problem, thereby proving that \textsc{BMA-GAI} is first-order optimal away from critical budget levels. Experiments on synthetic benchmarks and an LLM answer-verification task show that \textsc{BMA-GAI} allocates the verification budget more efficiently and certifies more high-quality candidates than competing methods.
Problem

Research questions and friction points this paper is trying to address.

Test-Time Scaling
Multi-Attribute Verification
Budgeted Good-Arm Identification
LLM Answer Verification
Innovation

Methods, ideas, or system contributions that make the work stand out.

Test-Time Scaling
Multi-Attribute Good-Arm Identification
Budgeted Verification
Anytime-Valid Certification
Cost-Aware Sampling
🔎 Similar Papers
No similar papers found.