Score
Design and implement mechanisms that use model-produced confidence or cached-distribution scores to decide whether to accept an output immediately or route it to further verification (e.g., debates or extra computation). This includes defining acceptance thresholds and gating policies, calibrating and evaluating confidence estimates, implementing early-accept/selective-gating logic to avoid unnecessary or error-amplifying verification, and bounding/tolerating rare false accepts.
The test-time scaling (TTS) field lacks a systematic survey, unified taxonomy, and principled analysis of verification methods. Method: We propose the first taxonomy framework for TTS verifiers, structured along three dimensions—verifier type, training paradigm (prompt-based guidance, discriminative/generative fine-tuning), and application mode—and integrate search-space exploration with candidate-output scoring for efficient inference optimization. We further provide a comprehensive survey of existing verification techniques and release an open-source verification resource repository. Contribution/Results: Our framework fills a critical gap in systematic TTS verification research, significantly improving the accuracy and reliability of large language model inference. It establishes a standardized foundation and reproducible benchmark for verifier design in TTS, enabling principled development and evaluation of verification mechanisms.
This paper studies the design of information verification mechanisms: how to probabilistically test agents’ reports to balance allocation efficiency and principal surplus. We propose a novel paradigm embedding statistical hypothesis testing into mechanism design, constructing a commitment mechanism with randomized verification—where each report type undergoes a binary test (pass/fail), and outcomes directly inform allocation and payment decisions. Innovatively, we reformulate the virtual value function and, under quasilinear preferences, derive the first closed-form solution for the optimal verification mechanism. Theoretically, we prove that higher verification accuracy strictly improves both allocation efficiency and the principal’s share of surplus, and that these two objectives are positively correlated. Our results establish a theoretically rigorous yet practically implementable foundation for credible information elicitation.
This study addresses the stagnation of enterprise AI initiatives in regulated financial institutions due to the absence of quantifiable evaluation criteria. Focusing on six document-intensive workflows, it systematically compares AI system performance across four model families and three tool configurations, distinguishing between demonstration and production environments. For the first time, it links deployment feasibility with human review rates. The authors propose a production-grade evaluation framework encompassing accuracy, reproducibility, traceability, and informative confidence, integrating multi-model comparison, confidence signals, source citation, and self-verification mechanisms. Experiments reveal that 56.1% of the 72 evaluated configurations meet production readiness thresholds. Incorporating source citation and confidence estimation reduces human review requirements to 49%, and adding self-verification further lowers this to 44%, albeit at the cost of reduced error tolerance.
Current large reasoning models fail to effectively model the relationship between confidence distributions and accuracy in multi-candidate answer selection, leading to insufficient reliability. This work proposes DistriVoting, a novel approach that explicitly decomposes confidence distributions into positive and negative components, models them via a Gaussian mixture model, and introduces a rejection filter to reduce distributional overlap. Additionally, it designs a SelfStepConf mechanism that dynamically adjusts the reasoning process based on step-level confidence to enhance distribution separation. By integrating a distribution-guided voting scheme with a dynamic reasoning strategy, the method significantly outperforms state-of-the-art approaches across 16 models and 5 benchmarks, substantially improving both answer selection accuracy and confidence calibration.
This work addresses the inefficiency and potential errors in inference services caused by redundant verification or blind correction. It proposes SevRA, a service-layer controller that, for the first time, formulates selective verification as a resource allocation problem at deployment time. Operating under the constraint of freezing the original large model’s outputs (e.g., Qwen3-4B), SevRA employs a recoverability-aware gating mechanism to trigger verification only when necessary. The approach requires no modification to the solver and instead trains a lightweight policy solely from log data, leveraging features of the attempted state to decide intervention. Experiments show that on MathFive, SevRA achieves 76.3% accuracy while reducing post-generation tokens by 26.8% and harmful flips by 54.5%; on GSM, it boosts accuracy to 94.5% by verifying only 3% of samples, saving 91.2% of verification overhead.
Existing generate-verify reasoning paradigms lack an explicit monitoring mechanism, preventing models from assessing task difficulty and self-confidence *prior* to generation—leading to the “prefix-dominance trap” and ~20% accuracy degradation. To address this, we propose the Monitor-Generate-Verify (MGV) framework, the first computationally grounded instantiation of Flavell’s and Nelson–Narens’ metacognitive theories in large language model reasoning. MGV introduces a pre-generation monitoring module—comprising explicit difficulty estimation and confidence calibration—coupled with a feedback-driven self-regulation mechanism that dynamically modulates the generation process. By fundamentally closing the monitoring gap inherent in conventional paradigms, MGV establishes a principled framework for diagnosing reasoning failures. It advances test-time reasoning architectures by enhancing both interpretability and robustness, offering a novel paradigm and concrete, actionable pathways for improvement.
This work addresses the performance degradation often observed in self-improving agents due to a misalignment between self-generated validation signals and actual deployment performance. To mitigate this issue, the authors propose the Sealed External Audit Loop (SEAL) mechanism, which preserves the agent’s ability to generate its own test cases while incorporating an external audit signal that the agent cannot manipulate. SEAL enforces a fixed audit loop, compares candidate policies, conducts sealed evaluations, and maintains state consistency to effectively block harmful updates. Experimental results across six models and three random seeds demonstrate that SEAL significantly outperforms unprotected baselines, robustly alleviating self-validation failure—particularly in capability-stratified agents.
This work addresses the prevalent issue in autonomous coding agents that prematurely declare lifecycle states—such as “DONE”—without verification during multi-step software tasks, often leading to erroneous progression. To mitigate this, the authors propose Proof-or-Stop, a model-agnostic and platform-neutral trusted control layer that strictly gates state transitions only when fresh, traceable, and mechanically verifiable evidence satisfies predefined conditions. Crucially, the approach treats agent outputs as claims pending validation rather than established facts, and explicitly distinguishes between the mere existence of review mechanisms and their role as gating criteria. Empirical evaluation demonstrates zero false “DONE” declarations across ten scenarios, successful resistance against 18 classes of tampering attacks with no false acceptances, and a reduction in hidden failure rates from 31/1800 to 2/1800 in ablation studies. Furthermore, 94.8% of issues in a corpus of 565 self-application narratives were resolved.
This work addresses the inefficiency of uniform computation allocation in existing language models during reasoning and the impracticality of deploying external-feedback-based verification mechanisms. The authors propose Self-Verification Refinement (SVR), a novel framework that leverages self-verification as an internal control signal to enable adaptive, unsupervised computation allocation at test time. SVR jointly optimizes answer correctness and confidence estimation through multi-round reinforcement learning, dynamically deciding whether to continue reasoning. It employs a GRPO algorithm trained with a composite reward combining correctness, calibration-aware self-verification, and stop-state incentives. Evaluated on seven mathematical reasoning benchmarks using Qwen3.5-2B, SVR achieves an average accuracy of 0.563 in just 2.99 reasoning rounds, significantly outperforming standard GRPO, multi-round baselines, and fixed-budget oracle-guided approaches.