🤖 AI Summary
This work addresses the limitations of existing test-time scaling methods, which struggle to dynamically allocate computational resources according to problem difficulty under a fixed budget and lack interpretability. The authors propose an adaptive sampling mechanism based on a lightweight fuzzy controller that jointly considers prompt complexity and model confidence to dynamically adjust the number of samples per query, enabling efficient and transparent inference. By integrating interpretable signals into a fuzzy control framework for the first time, the method significantly outperforms multiple baselines on question-answering and mathematical reasoning tasks: it reduces average sampling counts while maintaining accuracy close to that of full-budget controls, effectively balancing efficiency, performance, and decision interpretability.
📝 Abstract
Test-time scaling improves LLM reasoning by generating and aggregating multiple candidate answers, yet many pipelines use fixed per-query budgets that spend the same compute on easy and difficult prompts. These fixed budgets are also difficult to inspect because they do not explain why a given prompt receives a particular number of samples. We propose adaptive} test-time scaling with a lightweight fuzzy controller that maps interpretable signals, including estimated prompt complexity and model confidence, to a per-query sampling budget. The controller assigns fewer samples to easier or more confident prompts and more samples to harder or less certain prompts, making inference-time compute inspectable rather than fixed or opaque. We evaluate under a fair-alignment protocol with matched decoding settings and controlled answer selection, and compare against best-of-$N$, compute-aware scaling, and self-certainty-based baselines on question-answering and mathematical reasoning tasks. Across models and datasets, adaptive fuzzy control improves over several standard baselines and remains close to a selector-matched full-budget control while reducing the average number of samples. These findings suggest that interpretable adaptive sampling is a practical direction for more efficient test-time reasoning in large language models.