adaptive test-time scaling

Designs and evaluates inference-time mechanisms that dynamically allocate computation or model capacity per input by scaling steps, internal iterations, ensemble members, or precision according to model confidence signals (including verbalized confidence). The work builds confidence-thresholding and routing policies and evaluation metrics to reduce average inference cost while preserving accuracy through per-example, confidence-guided scaling.

adaptivetest-timescaling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.01
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$201K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search

Apr 11, 2026
AR
Austin R. Ellis-Mohr
🏛️ University of Illinois Urbana-Champaign

LLM inference cost is increasingly becoming a critical resource bottleneck; existing optimality analyses often decouple training from inference and neglect dynamic trade-offs in strategy selection. This paper proposes Directed Stochastic Skill Search (DS3), a framework that models inference as stochastic traversal over a skill graph, establishing the first unified theoretical model for joint training–inference optimization. Leveraging a tripartite graph structure and stochastic process theory, we derive closed-form expressions for computational cost and accuracy of inference strategies—including Chain-of-Thought (CoT) and Tree-of-Thought (ToT)—and rigorously characterize emergence conditions for mechanisms such as Bag-of-Natural-proofs (BoN) and majority voting from first principles. Our theory reproduces scaling phenomena—e.g., linear accuracy growth under logarithmic compute—and yields a critical threshold under which small models can surpass large ones via efficient inference. The results provide quantitative design principles for low-cost, high-reliability LLM inference.

Analyzing task success and compute cost across inference methodsExploring efficient inference strategies beyond fixed compute-optimalityOptimizing inference costs for large language models (LLMs)

This work addresses the lack of a unified framework in existing test-time scaling methods, which hinders fair comparison of inference algorithms in terms of computational budgets, evaluation metrics, and reproducibility. We propose a formal budgeted inference framework grounded in prefix trees, systematically distinguishing three inference architectures: single-trajectory sequential expansion, leaf-node aggregation, and prefix-level search. For the first time, we establish a three-dimensional analytical framework encompassing structural taxonomy, end-to-end evaluation protocols, and reproducibility standards. By formalizing reasoning-tree modeling, introducing multi-dimensional evaluation profiles, and aligning computational and uncertainty reporting mechanisms, we integrate open-source model ecosystems and validate our approach on benchmarks spanning general knowledge, symbolic reasoning, and competition mathematics. We release over two billion complete reasoning trajectories alongside verifiers and token-level signals.

evaluationinference protocolsreasoning LLMs

This work proposes a verifier-guided adaptive inference framework that overcomes the inefficiencies of static computation allocation in conventional test-time reasoning. By modeling inference as an iterative process of trajectory generation and selection, the method dynamically plans, selects tools, and adjusts computational strategies at each step, all under the unified guidance of a Process Reward Model (PRM). This approach achieves, for the first time, fine-grained, cross-iteration adaptive computation allocation based on PRM signals, transcending the limitations of fixed sampling and post-hoc reranking. Evaluated on challenging benchmarks—including MATH-500, AIME24, and AMO-Bench—the framework significantly outperforms existing test-time scaling methods, delivering higher accuracy while reducing wasteful generations and tool invocation overhead.

adaptive allocationcompute efficiencyreasoning trajectories

Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Aug 01, 2024
YW
Yangzhen Wu
🏛️ Tsinghua University | Carnegie Mellon University

This work addresses computational efficiency optimization during large language model (LLM) inference, systematically investigating the trade-off between model scale and generated token count. We propose a novel tree-search inference algorithm and comprehensively evaluate cost-performance Pareto frontiers under varying compute budgets—employing greedy search, best-of-n sampling, and weighted voting—across Llemma models of multiple scales (7B–34B). Experimental results demonstrate that scaling inference compute yields substantially greater gains than scaling model parameters. Notably, Llemma-7B equipped with our tree-search algorithm consistently outperforms Llemma-34B and all baseline methods on the MATH benchmark, achieving Pareto-optimal “small-model + strong-inference” performance. This work provides both a scalable algorithmic framework and empirical evidence for efficient LLM deployment.

Demonstrates smaller models with advanced algorithms outperform larger models.Explores optimal inference configurations for large language models.Investigates trade-offs between model size and token generation strategies.

Existing sequential scaling methods predominantly rely on heuristic strategies, lacking theoretical guarantees that limit both performance and interpretability. This work pioneers a formal modeling of sequential scaling as a two-state Markov process, from which we derive sufficient conditions for accuracy improvement and establish provable upper and lower bounds on performance. Building upon this theoretical foundation, we develop a closed-form optimization solution that enables principle-driven inference scheduling. Evaluated across three prominent large language models, five benchmark datasets, and over twenty experimental configurations, our approach consistently outperforms existing parallel and sequential scaling strategies, achieving significant gains in both inference efficiency and accuracy.

inference-time scalinglarge language modelsMarkov process

Latest Papers

What's happening recently
View more

This study addresses the challenge of balancing performance and efficiency in large language model inference under constrained computational resources. Through systematic evaluation across model scales and configurations on the MMLU-Pro and BBH benchmarks, it investigates reasoning augmentation strategies—including self-consistency, self-refinement, multi-agent debate, and agent ensembling. Large-scale multi-configuration experiments reveal that multi-agent approaches consistently enhance performance on highly complex tasks. Under identical compute budgets, debate and agent ensembling outperform self-consistency by 1.3% and 2.7% in accuracy, respectively. With a 20× Chain-of-Thought (CoT) budget, reasoning augmentation boosts MMLU-Pro accuracy by up to 7.1%. The work further establishes practical guidelines for efficient agent ensembling and leverages Pareto frontiers to inform optimal strategy selection.

compute efficiencyinference scalingmulti-agent reasoning

This study systematically evaluates the effectiveness of inference-time optimization strategies for large language models on mathematical reasoning tasks, with a focus on International Mathematical Olympiad (IMO)-level problems. The central hypothesis is that reducing error correlation across multiple solution attempts can improve overall accuracy. We present the first large-scale empirical assessment of techniques such as diversified prompting, high-temperature sampling, and multi-model ensembling during the AIMO 3 competition. Our findings indicate that none of these interventions yield significant performance gains: high-temperature sampling alone suffices to decorrelate errors, while weakened prompts actually degrade single-attempt accuracy. Moreover, differences in base model capabilities exert an order-of-magnitude greater influence on performance than any inference-time optimization strategy examined.

diverse promptingerror correlationinference-time optimization

This study addresses the performance bottlenecks of resource-constrained local agents under stringent hardware limitations by systematically investigating four dimensions of inference-time scaling: context, time, structure, and parallelism. Experiments on the OSWorld benchmark using Qwen3-VL-8B/30B-A3B, UI-TARS-1.5-7B, and OpenCUA-7B models reveal that context extension enhances trajectory stability but exhibits diminishing returns, temporal extension alleviates stuttering yet yields limited gains in task success rates, and parallel execution reduces structural overhead at the cost of high computational expense. The work uncovers, for the first time, the law of diminishing returns in inference scaling and a shift in failure modes, leading to a novel paradigm that integrates selective computation allocation with failure-aware control—offering both theoretical grounding and practical pathways for designing efficient local agents.

compute tradeoffscomputer-use agentsfailure modes

This study addresses the limitations of dense large language models in long-chain reasoning, where fragmented KV caches and parallelization inefficiencies undermine traditional prefill extension strategies. Through systematic evaluation of dense and Mixture-of-Experts models ranging from 8B to 671B parameters on GPU clusters, the work uncovers critical performance bottlenecks: a sharp drop in data parallelism efficiency due to cache fragmentation, a nonlinear scaling inflection point in tensor parallelism around 32B parameters, and fundamental differences between sparse and dense architectures in interconnect bandwidth utilization and routing latency. Guided by extensive empirical analysis, the authors propose an architecture decision framework tailored to the “inference cliff” phenomenon, establishing design principles for next-generation LLM inference infrastructure that substantially improve resource utilization and throughput efficiency.

capacity-bound regimeinference scalingKV-cache fragmentation

Hot Scholars

GL

Guilin Liu

Research Scientist, NVIDIA
Computer VisionDeep LearningGenerative Models
YS

Yuxuan Song

Tsinghua University
Deep Generative ModelsLLM4Science,
SB

Sebastian Baltes

University of Bayreuth
software engineeringempirical software engineering
ZL

Ziquan Liu

Assistant Professor, Queen Mary University of London
machine learning