Risk-Controlled Selective LLM Answering by Pricing Label-Free Checks

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of relying solely on verifier ranking to control error rates in large language model (LLM) serving. To this end, it proposes PriceCheck, a framework that introduces a price-based coverage prediction mechanism integrating verifier rankings with the cost of unlabeled consistency checks. Through price-combined decision rules, PriceCheck dynamically guides check scheduling and termination strategies, enabling precise control over selective risk. Experimental results demonstrate that on mathematical tasks, PriceCheck serves 76.1% of answers on average while maintaining a selective risk below 1.5%, significantly outperforming existing baseline methods. This work establishes an efficient cost-risk control paradigm for the reliable deployment of LLMs.
📝 Abstract
Serving an answer from a large language model requires deciding when to abstain, yet a verifier's ranking accuracy alone does not determine the error rate among served answers. We introduce PriceCheck, which builds a compact family of decision rules from label-free checks such as re-solving a problem. Each check has a price: its agreement rates on correct and incorrect answers and its cost per run. Prices fitted on a small, class-enriched labelled set compose into predictions of a schedule's coverage and cost, guiding which checks to run and when to stop. A calibration test then selects a schedule at a stated selective-risk target. In mathematics, the selected schedules serve 76.1% of answers on average and keep held-out selective risk below 1.5% on all 15 splits. Under the shared testing protocol, PriceCheck serves more answers at that target than reward models, a prompted judge, the generator's confidence and a trained correctness classifier. At matched coverage, it keeps the fewest wrong answers among these scorers. Across 118 diagnostic schedules, price-based coverage predictions have a rank correlation of 0.97 with observed coverage. These results show that choosing how checks are combined and stopped matters alongside how well a verifier ranks answers. Code is available at https://github.com/js-lee-AI/PriceCheck.
Problem

Research questions and friction points this paper is trying to address.

selective answering
risk control
large language models
abstention
label-free verification
Innovation

Methods, ideas, or system contributions that make the work stand out.

Selective Answering
Label-Free Checks
Risk Control
Decision Rules
Coverage Prediction
🔎 Similar Papers
No similar papers found.