Online Verification of Language Model Responses Under Cost Constraints

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge in LLM output verification where a fixed weak verifier struggles to adapt to dynamic queries, resulting in an imbalance between cost and accuracy. To this end, this work proposes the OMVV algorithm, which maintains a pool of candidate weak verifiers and employs online learning with an exponential weighting strategy to achieve adaptive routing decisions. Theoretically, the algorithm provides distribution-free, finite-time error rate guarantees while attaining sublinear regret under cost and consistency constraints. Empirically, evaluations on reasoning benchmarks demonstrate that the proposed method achieves superior accuracy compared to any single fixed verifier at a lower verification cost.
📝 Abstract
As large language models are increasingly deployed for multi-step reasoning, verifying the correctness of their outputs has become essential for maintaining reliability at scale. Verifying the correctness of large language model outputs is often done by querying a costly ground-truth oracle, which is impractical to invoke at every step in an online setting. Prior work addresses this by querying a single weak verifier on every step, and using its score to decide whether the costly strong verifier needs to be queried as well, reserving strong verification for only a small fraction of the steps. However, a single fixed weak verifier may not perform consistently well as the subject matter or difficulty of incoming queries changes over time, and committing to one in advance risks either overly costly or inaccurate verification. We introduce OMVV (Online Multi-Verifier Verification), an algorithm that maintains a pool of $K$ candidate weak verifiers with differing cost and verification performance, and adaptively routes each round's decision to a verifier selected via an online score combiner and an exponential-weights routing policy. OMVV provides a distribution-free, finite-time guarantee on false-accept and false-reject rates across the full pool of verifiers, and further achieves sublinear regret against the best fixed verifier in hindsight under a combined cost and consistency objective. Experiments on reasoning dataset benchmarks show that OMVV achieves higher accuracy at lower verification cost than any single fixed verifier, across a range of operating budgets.
Problem

Research questions and friction points this paper is trying to address.

large language models
online verification
cost constraints
multi-verifier
multi-step reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Online Multi-Verifier Verification
Exponential-Weights Routing
Cost-Constrained Verification
Sublinear Regret
Large Language Models
💼 Related Jobs
No related jobs found.
E
Erfan Hajihashemi
Department of Electrical Engineering & Computer Science, University of California, Irvine
Yanning Shen
Yanning Shen
University of California, Irvine
Trustworthy ML/AILearning over GraphsOnline Learning