Taming Speculative Search for Test-Time Scaling in LLM Serving

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the search space explosion and frequent fine-grained verification overhead induced by speculative execution during test-time scaling of large language models. To this end, we propose SpecScale, a system that synergistically integrates three mechanisms: early pruning to eliminate invalid reasoning paths, computation deduplication to remove redundant exploration, and deferred verification to reduce validation frequency. Together, these components effectively balance inference efficiency against computational cost. Experimental results demonstrate that SpecScale significantly outperforms existing non-speculative and speculative approaches on benchmarks such as MATH. By substantially improving throughput and reducing latency while preserving generation quality, this work establishes a new paradigm for efficient test-time scaling.
📝 Abstract
Test-time scaling has recently emerged as a powerful approach for improving LLM reasoning by allocating additional computation during inference, substantially enhancing accuracy on challenging tasks such as mathematics and coding. To accelerate the exploration of reasoning paths, recent studies proposed speculative execution. However, we show that supporting speculative execution poses two unique challenges for LLM serving systems: (1) an explosion in the search space of candidate paths and (2) frequent, fine-grained verification tasks for candidates. To address these challenges, this paper proposes SpecScale, a serving system for efficient speculative execution. We introduce three techniques to reconcile the trade-off between latency and computational overhead: (1) early pruning of low-quality candidate paths, (2) deduplicating computation across redundant candidate paths, and (3) deferring fine-grained verification tasks. We evaluate SpecScale on challenging reasoning benchmarks, including MATH and Olympiad. Our results show that SpecScale significantly outperforms both non-speculative and recent speculative approaches, delivering substantial improvements in throughput and latency while preserving answer quality.
Problem

Research questions and friction points this paper is trying to address.

Test-Time Scaling
Speculative Execution
LLM Serving
Search Space Explosion
Fine-grained Verification
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speculative Execution
Test-Time Scaling
Early Pruning
Computation Deduplication
LLM Serving