Score
Designs and evaluates inference-time mechanisms that dynamically allocate computation or model capacity per input by scaling steps, internal iterations, ensemble members, or precision according to model confidence signals (including verbalized confidence). The work builds confidence-thresholding and routing policies and evaluation metrics to reduce average inference cost while preserving accuracy through per-example, confidence-guided scaling.
Large language models (LLMs) suffer from inefficient inference—fixed computational budgets poorly align with varying task complexity, causing over-computation on simple tasks and under-computation on complex ones. Method: This paper presents a systematic survey of adaptive and controllable test-time computation (TTC) strategies, proposing a two-level taxonomy: L1 (controlled inference under fixed budget) and L2 (dynamic resource allocation). It innovatively integrates dynamic scaling, confidence-guided early exiting, and hybrid inference modes to jointly optimize token efficiency and performance. Contribution/Results: Empirical evaluation across multiple benchmarks on mainstream closed-source LLMs establishes the first quantitative characterization of the trade-off between inference efficacy and computational cost. The work delivers both a theoretical framework and an empirical benchmark for efficient, user-constrained, and resource-adaptive LLM inference—emphasizing practicality, scalability, and responsiveness to user-specified constraints.
LLM inference cost is increasingly becoming a critical resource bottleneck; existing optimality analyses often decouple training from inference and neglect dynamic trade-offs in strategy selection. This paper proposes Directed Stochastic Skill Search (DS3), a framework that models inference as stochastic traversal over a skill graph, establishing the first unified theoretical model for joint training–inference optimization. Leveraging a tripartite graph structure and stochastic process theory, we derive closed-form expressions for computational cost and accuracy of inference strategies—including Chain-of-Thought (CoT) and Tree-of-Thought (ToT)—and rigorously characterize emergence conditions for mechanisms such as Bag-of-Natural-proofs (BoN) and majority voting from first principles. Our theory reproduces scaling phenomena—e.g., linear accuracy growth under logarithmic compute—and yields a critical threshold under which small models can surpass large ones via efficient inference. The results provide quantitative design principles for low-cost, high-reliability LLM inference.
This work addresses the lack of a unified framework in existing test-time scaling methods, which hinders fair comparison of inference algorithms in terms of computational budgets, evaluation metrics, and reproducibility. We propose a formal budgeted inference framework grounded in prefix trees, systematically distinguishing three inference architectures: single-trajectory sequential expansion, leaf-node aggregation, and prefix-level search. For the first time, we establish a three-dimensional analytical framework encompassing structural taxonomy, end-to-end evaluation protocols, and reproducibility standards. By formalizing reasoning-tree modeling, introducing multi-dimensional evaluation profiles, and aligning computational and uncertainty reporting mechanisms, we integrate open-source model ecosystems and validate our approach on benchmarks spanning general knowledge, symbolic reasoning, and competition mathematics. We release over two billion complete reasoning trajectories alongside verifiers and token-level signals.
This work proposes a verifier-guided adaptive inference framework that overcomes the inefficiencies of static computation allocation in conventional test-time reasoning. By modeling inference as an iterative process of trajectory generation and selection, the method dynamically plans, selects tools, and adjusts computational strategies at each step, all under the unified guidance of a Process Reward Model (PRM). This approach achieves, for the first time, fine-grained, cross-iteration adaptive computation allocation based on PRM signals, transcending the limitations of fixed sampling and post-hoc reranking. Evaluated on challenging benchmarks—including MATH-500, AIME24, and AMO-Bench—the framework significantly outperforms existing test-time scaling methods, delivering higher accuracy while reducing wasteful generations and tool invocation overhead.
This work addresses computational efficiency optimization during large language model (LLM) inference, systematically investigating the trade-off between model scale and generated token count. We propose a novel tree-search inference algorithm and comprehensively evaluate cost-performance Pareto frontiers under varying compute budgets—employing greedy search, best-of-n sampling, and weighted voting—across Llemma models of multiple scales (7B–34B). Experimental results demonstrate that scaling inference compute yields substantially greater gains than scaling model parameters. Notably, Llemma-7B equipped with our tree-search algorithm consistently outperforms Llemma-34B and all baseline methods on the MATH benchmark, achieving Pareto-optimal “small-model + strong-inference” performance. This work provides both a scalable algorithmic framework and empirical evidence for efficient LLM deployment.
Existing sequential scaling methods predominantly rely on heuristic strategies, lacking theoretical guarantees that limit both performance and interpretability. This work pioneers a formal modeling of sequential scaling as a two-state Markov process, from which we derive sufficient conditions for accuracy improvement and establish provable upper and lower bounds on performance. Building upon this theoretical foundation, we develop a closed-form optimization solution that enables principle-driven inference scheduling. Evaluated across three prominent large language models, five benchmark datasets, and over twenty experimental configurations, our approach consistently outperforms existing parallel and sequential scaling strategies, achieving significant gains in both inference efficiency and accuracy.
This study addresses the challenge of balancing performance and efficiency in large language model inference under constrained computational resources. Through systematic evaluation across model scales and configurations on the MMLU-Pro and BBH benchmarks, it investigates reasoning augmentation strategies—including self-consistency, self-refinement, multi-agent debate, and agent ensembling. Large-scale multi-configuration experiments reveal that multi-agent approaches consistently enhance performance on highly complex tasks. Under identical compute budgets, debate and agent ensembling outperform self-consistency by 1.3% and 2.7% in accuracy, respectively. With a 20× Chain-of-Thought (CoT) budget, reasoning augmentation boosts MMLU-Pro accuracy by up to 7.1%. The work further establishes practical guidelines for efficient agent ensembling and leverages Pareto frontiers to inform optimal strategy selection.
This study systematically evaluates the effectiveness of inference-time optimization strategies for large language models on mathematical reasoning tasks, with a focus on International Mathematical Olympiad (IMO)-level problems. The central hypothesis is that reducing error correlation across multiple solution attempts can improve overall accuracy. We present the first large-scale empirical assessment of techniques such as diversified prompting, high-temperature sampling, and multi-model ensembling during the AIMO 3 competition. Our findings indicate that none of these interventions yield significant performance gains: high-temperature sampling alone suffices to decorrelate errors, while weakened prompts actually degrade single-attempt accuracy. Moreover, differences in base model capabilities exert an order-of-magnitude greater influence on performance than any inference-time optimization strategy examined.
This study addresses the performance bottlenecks of resource-constrained local agents under stringent hardware limitations by systematically investigating four dimensions of inference-time scaling: context, time, structure, and parallelism. Experiments on the OSWorld benchmark using Qwen3-VL-8B/30B-A3B, UI-TARS-1.5-7B, and OpenCUA-7B models reveal that context extension enhances trajectory stability but exhibits diminishing returns, temporal extension alleviates stuttering yet yields limited gains in task success rates, and parallel execution reduces structural overhead at the cost of high computational expense. The work uncovers, for the first time, the law of diminishing returns in inference scaling and a shift in failure modes, leading to a novel paradigm that integrates selective computation allocation with failure-aware control—offering both theoretical grounding and practical pathways for designing efficient local agents.
This study addresses the limitations of dense large language models in long-chain reasoning, where fragmented KV caches and parallelization inefficiencies undermine traditional prefill extension strategies. Through systematic evaluation of dense and Mixture-of-Experts models ranging from 8B to 671B parameters on GPU clusters, the work uncovers critical performance bottlenecks: a sharp drop in data parallelism efficiency due to cache fragmentation, a nonlinear scaling inflection point in tensor parallelism around 32B parameters, and fundamental differences between sparse and dense architectures in interconnect bandwidth utilization and routing latency. Guided by extensive empirical analysis, the authors propose an architecture decision framework tailored to the “inference cliff” phenomenon, establishing design principles for next-generation LLM inference infrastructure that substantially improve resource utilization and throughput efficiency.