🤖 AI Summary
This study addresses the challenge in large language model (LLM) voting whereby identifying correct answers under a fixed budget does not effectively translate into final decisions. Assuming independent and identically distributed responses, this work employs state-space analysis to quantify the recoverability threshold from discovery to decision-making. It derives exact finite-horizon endpoint probabilities and gold-label-free locking certificates, revealing the transition window between candidate set expansion and winner contraction, as well as the impact of erroneous merging on precision. Evaluated on the Word16 benchmark, the proposed approach improves accuracy by 21.1% and reduces API invocation overhead by 28–30% while preserving output consistency.
📝 Abstract
Voting over multiple LLM responses is a common primitive in test-time scaling and ensemble inference. Collecting more responses can expand the candidate pool and increase the chance that a correct answer is discovered. Under a fixed call budget, a discovered answer still needs to accumulate enough support within the remaining calls to become the final plurality winner, creating a discovery-to-decision gap. In this work, we characterize this gap through the realized vote state and remaining call budget. We derive a sharp recoverability threshold and show that, as sampling proceeds, the observed candidate set can only expand while the set of reachable endpoint winners can only contract, inducing a candidate-level conversion window. Under a specified iid response law, the same state yields exact finite-horizon endpoint probabilities. We further show that merging wrong-answer identities preserves single-call correctness and cannot improve plurality accuracy, and that the effect of redistributing wrong-answer probability depends on the realized vote state. Singleton reachability yields a gold-free exact locking certificate. For a known answer universe, its first trigger is the earliest prefix at which all admissible continuations yield the same fixed-budget output. Empirically, most discovered-but-unselected correct answers lose reachability only after discovery. In a controlled Word16 study, input permutation improves raw-plurality accuracy by 21.1 points with essentially unchanged single-call correctness. Exact locking saves 28-30% of calls at a 16-call budget while preserving every fixed-budget output.