Strategic Self-Consistency

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability wherein large language model (LLM) providers can exploit self-consistency majority voting mechanisms to fraudulently overcharge users by generating redundant reasoning paths. Through multi-path generation and statistical distribution analyses conducted on Llama, Qwen, and DeepSeek-R1, this work reveals that additional paths exhibit heavy-tailed distribution characteristics. Building upon this finding, an efficient path re-ranking algorithm is proposed that renders redundant paths formally necessary without increasing computational costs, thereby evading auditing detection. The results demonstrate that even under stringent auditing conditions with a false positive rate below 0.1, the system retains substantial capacity for overcharging. This research exposes deep-seated security vulnerabilities in current LLM billing and auditing mechanisms, highlighting the urgent need for more robust verification frameworks.
📝 Abstract
Self-consistency has become a popular technique for enhancing the reasoning abilities of large language models by generating multiple reasoning paths and selecting the final answer through a majority vote. However, because model providers typically charge users in proportion to the number of reasoning paths generated, they have a financial incentive to artificially increase the path count. In this work, we show that an unfaithful provider can exploit this incentive using a simple, efficient algorithm while avoiding detection by an auditor: by generating and strategically reordering additional reasoning paths, the algorithm makes every path appear necessary to reach the majority. To validate our algorithm, we conduct experiments with multiple instruct models from the Llama and Qwen families, as well as reasoning models distilled from DeepSeek-R1, on benchmark datasets spanning mathematics, science, and question answering. Our results suggest that the distribution of additional reasoning paths generated by our algorithm is heavy-tailed and that substantial capacity to overcharge remains even under the best possible audit designed to keep the false-positive rate below $α= 0.1$.
Problem

Research questions and friction points this paper is trying to address.

Self-Consistency
Large Language Models
Overcharging
Reasoning Paths
Audit Evasion
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Consistency
Strategic Reordering
Large Language Models
Audit Evasion
Heavy-tailed Distribution
🔎 Similar Papers
No similar papers found.
T
Tori Qiu
Carnegie Mellon University, Pittsburgh, USA
Ander Artola Velasco
Ander Artola Velasco
PhD candidate, Max Planck Institute for Software Systems
StatisticsMachine Learning
M
Manuel Gomez-Rodriguez
Max Planck Institute for Software Systems, Kaiserslautern, Germany