A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search

📅 2026-04-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
LLM inference cost is increasingly becoming a critical resource bottleneck; existing optimality analyses often decouple training from inference and neglect dynamic trade-offs in strategy selection. This paper proposes Directed Stochastic Skill Search (DS3), a framework that models inference as stochastic traversal over a skill graph, establishing the first unified theoretical model for joint training–inference optimization. Leveraging a tripartite graph structure and stochastic process theory, we derive closed-form expressions for computational cost and accuracy of inference strategies—including Chain-of-Thought (CoT) and Tree-of-Thought (ToT)—and rigorously characterize emergence conditions for mechanisms such as Bag-of-Natural-proofs (BoN) and majority voting from first principles. Our theory reproduces scaling phenomena—e.g., linear accuracy growth under logarithmic compute—and yields a critical threshold under which small models can surpass large ones via efficient inference. The results provide quantitative design principles for low-cost, high-reliability LLM inference.

Technology Category

Search and Optimization: Learning to SearchReasoning under Uncertainty: Stochastic OptimizationConstraint Satisfaction and Optimization: Satisfiability Modulo Theories

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Large language models (LLMs) demand considerable computational, energy, and financial resources during both training and deployment. While scaling laws for training have guided much of the field's recent progress, inference costs now represent a significant and growing component of the overall resource burden, particularly for reasoning-focused models. Existing characterizations of compute-optimality that consider model size, dataset size, and inference tokens in isolation or in fixed combinations risk overlooking more efficient operating points. We introduce directed stochastic skill search (DS3), a general framework that represents inference as stochastic traversal over a learned skill graph. From a simplified yet expressive instantiation, we derive closed-form expressions for task success and compute cost across a wide range of inference strategies -- including chain-of-thought (CoT) and tree-of-thought (ToT) -- enabling comparative analysis as a function of task difficulty and model capability. To that end, we extend a prior first-principles tripartite graph framework of LLM training to incorporate inference, and separately bridge DS3 with empirical methods that characterize LLM scaling behavior. We theoretically recover empirically observed patterns, including: linear accuracy scaling with logarithmic compute; variation in preferred inference strategies as a function of task difficulty and model capability; emergent behavior elicited by reasoning even when performance plateaus under parameter scaling; and both best-of-N (BoN) and majority voting behavior captured within a unified analytical framework. By explicitly characterizing training-inference interdependencies, our framework deepens theoretical understanding and supports principled algorithmic design and resource allocation.
Problem

Research questions and friction points this paper is trying to address.

Optimizing inference costs for large language models (LLMs)
Exploring efficient inference strategies beyond fixed compute-optimality
Analyzing task success and compute cost across inference methods
Innovation

Methods, ideas, or system contributions that make the work stand out.

Directed stochastic skill search for inference optimization
Closed-form expressions for task success and cost
Tripartite graph framework integrating training and inference