Score
Designs and implements systems that predict and launch likely external tool or API calls in parallel with model decoding by creating speculative branches or 'forks' whose executions run ahead of confirmed reasoning. It also builds the control logic to overlap tool execution with token generation, accept or merge speculative results when the final reasoning validates them, and revert to serial execution or discard speculative outputs when predictions are incorrect.
Existing speculative decoding (SD) suffers from substantial latency due to sequential execution of the draft and target models. This paper proposes SpecBranch, the first framework to achieve parallelization in speculative decoding—inspired by processor branch prediction. It introduces a rollback-aware branching mechanism and jointly optimizes adaptive draft length with hybrid confidence modeling, integrating implicit draft-model confidence and explicit target-model feature reuse. The method comprises four core components: branch-parallel decoding, hybrid draft generation, rollback-aware scheduling, and dynamic length control. Experiments demonstrate that SpecBranch achieves 1.8×–4.5× speedup over autoregressive decoding. Moreover, for low-alignment models, it reduces rollback tokens by 50%, significantly improving inference efficiency and deployment practicality.
This work addresses the significant latency incurred by sequential tool invocation in large language model (LLM) agents, which leads to GPU underutilization. The authors propose a training-free speculative execution mechanism that forks lightweight probes early during generation to predict subsequent tool calls and executes them in parallel, thereby hiding waiting time. This approach achieves the first training-free speculation without relying on auxiliary models, historical trajectories, or static workflow graphs. By integrating confidence-based filtering and partial token acceptance, it maintains or improves task accuracy while enhancing efficiency. The system is compatible with standard APIs and orthogonal to token-level speculative decoding. Experiments show an 18% reduction in P95 latency on the GAIA benchmark using Qwen3-32B, with consistent gains across models ranging from 4B to 32B parameters and accuracy matching or exceeding the baseline.
This work addresses the high latency of multi-step tool calling, which severely hinders the deployment of large language models in real-time services. To this end, it introduces a training-free, plug-and-play acceleration method that integrates structured tool-calling patterns and retrieval-augmented mechanisms into a speculative decoding framework for the first time. The approach employs a finite-state machine to alternately fill pattern tokens and speculatively generate variable fields, while leveraging vector retrieval to reuse historical tool-call records as drafts, substantially improving generation efficiency. Experimental results demonstrate that the proposed method achieves up to 4.2× inference speedup across multiple benchmarks, significantly outperforming existing training-free speculative decoding strategies.
Current LLM agents are constrained by a serial “LLM–tool” execution loop, resulting in significant latency. This work proposes PASTE, the first approach to enable speculative parallel execution of tool calls by exploiting stable control-flow patterns and predictable data dependencies inherent in tasks. PASTE employs a pattern-aware speculation mechanism, explicit control-flow modeling, and dynamic dependency analysis to schedule tool invocations in parallel, effectively masking execution latency. Experimental results demonstrate that PASTE reduces average task completion time by 48.5% compared to state-of-the-art methods and achieves a 1.8× improvement in tool throughput, substantially overcoming the bottleneck imposed by sequential execution.
Standard speculative decoding suffers from limited inference acceleration due to the strict serial dependency between draft generation and verification. This work proposes MineDraft, a novel batch-parallel speculative decoding framework that introduces a dual-batch pipelined scheduling mechanism, overlapping draft generation for one batch of requests with the verification phase of another. Furthermore, MineDraft incorporates a cooperative verification protocol between the draft and target models to ensure output correctness while significantly improving hardware utilization. Integrated into the vLLM system, the proposed approach reduces end-to-end latency by up to 39% and achieves a throughput improvement of up to 75% compared to standard speculative decoding.
Symbolic execution often struggles to adequately explore program paths due to resource constraints. To address this limitation, this work proposes Agolic, a novel system that introduces agent-based planning into the symbolic execution workflow. Without altering the underlying exploration logic, Agolic dynamically configures multiple rounds of bounded symbolic execution through cross-round, high-level reasoning. The approach synergistically integrates large language models, source code analysis, coverage replay, and goal-directed strategies to substantially enhance path coverage. Experimental results demonstrate that Agolic achieves, on average, more than three times the branch coverage of continuous symbolic execution across several C/C++ programs and uncovers previously unexplored branches in six out of seven benchmarks—branches missed by a combination of fuzzing and compiler-assisted concrete execution.
This work addresses the high latency and lack of parallelism in existing large language models that generate tool calls token-by-token, preventing concurrent prediction of functions and their parameters. To overcome this, the authors propose a lightweight sidecar model that, upon request arrival, simultaneously predicts both the function choice and all parameter slots in parallel. The predictions are then seamlessly integrated into the main model’s decoding process via non-blocking semantic injection. This approach achieves, for the first time, request-level out-of-order semantic speculation without requiring retraining of the drafter for each target model, offering strong generality and multi-model deployment capability. Evaluated with a LoRA-finetuned Qwen3-0.6B sidecar model combined with split-GPU pipeline parallelism, the method yields an average 3.89× speedup across 21 target-benchmark combinations—significantly outperforming ToolSpec (2.95×) and existing learned drafters by 34.1%.
Traditional program analysis relies on control flow graphs and separate forward or backward data-flow analyses, resulting in complex structures that are difficult to formally verify. This work proposes a forward-directed formal method based on prophecy and history variables, tightly integrating program analysis into operational semantics and establishing correctness and optimality of transformations via subset-inclusion constraints. We present the first machine-verified framework supporting prophecy and history variables, eliminating the need for explicit control flow graphs, abstraction/concretization functions, and Galois connections. By extending the operational semantics of a domain-specific language and employing forward simulation, we implement formally verified program transformations within the Nexis compiler. This approach yields the first machine-checked proofs of both correctness and optimality for dead code elimination and lazy code motion optimizations.
This work addresses the limitation of existing probing methods, which struggle to identify reusable internal causal interfaces in language models that support diverse future computations due to their reliance solely on current outputs. The authors propose a label-free framework for discovering such causal interfaces by introducing a “forked futures” mechanism: after a shared prefix, multiple divergent future trajectories are sampled to construct a causal quotient space based on comparisons of response distributions. Four interface types—including Shared—are defined and competitively evaluated using a preordered causal description length, with structural selection guided by a fidelity constraint on future signatures. Experiments demonstrate that the method reduces description length by 0.216 and 0.294 nats on Qwen2.5-1.5B and Llama-3-8B, respectively; it successfully recovers 14 out of 16 model architectures in blind tests and achieves an API alignment path mediation effect of 0.749.