Adaptive Self-Consistency: From Black-Box Sampling to Distribution-Valued Feedback

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the low sampling efficiency of black-box methods in large language model inference by introducing gray-box distributional information to optimize stopping strategies for the first time. We formulate efficient inference as a sequential pattern recognition problem with distributional observations, integrating Bayesian inference with gambling-based stopping rules to propose an adaptive stopping algorithm, ASC-D. Theoretically, we prove that the stopping rate under the gray-box setting is superior to or at least matches the theoretical limit of black-box approaches. Empirically, our method reduces the number of sampled trajectories by 46.4%–95.6% on MMLU-Redux and significantly improves the certificate rate under fixed computational budgets.
📝 Abstract
Self-consistency samples many reasoning trajectories and aggregates their final answers, treating the LLM as a black box that returns one answer per trajectory. Yet the final answer of each trajectory is sampled from a softmax vector that is available from the model's log-probabilities. We refer to this as the grey-box setting in which each trajectory reveals this answer distribution rather than a single draw from it. We formulate efficient inference in this setting as sequential mode identification with distribution-valued observations: sample trajectories one at a time and stop as soon as the LLM's modal answer is identified at a prescribed confidence level. We characterize the asymptotic stopping rate of mode identification with distribution-valued observations exactly and show that it is never worse than the black-box rate. We then propose the ASC-D algorithm, a betting stopping rule that attains this asymptotic stopping rate. On MMLU-Redux, ASC-D uses $46.4$--$95.6\%$ fewer trajectories than answer-only adaptive self-consistency baselines and achieves the highest fixed-budget correct-certification rate across three open-source models.
Problem

Research questions and friction points this paper is trying to address.

self-consistency
large language models
mode identification
distribution-valued observations
sequential inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive Self-Consistency
Grey-box Setting
Distribution-valued Feedback
Sequential Mode Identification
Betting Stopping Rule
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.