SchemaFill: Efficient LLM Tool Calling via Slot-Parallel Speculative Decoding

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Tool calling in large language models relies on autoregressive decoding, resulting in substantial generation latency in multi-parameter scenarios. This work proposes slot-parallel speculative decoding, a method that concurrently generates candidate values across fields and calls without prior knowledge of the invocation sequence. By integrating a structured output verification mechanism, the target model ensures generation accuracy. Evaluated on the Glaive and BFCL benchmarks, the proposed approach achieves up to a 4.05× improvement in end-to-end throughput compared to standard autoregressive decoding, enabling efficient and reliable tool calling for large language models.
📝 Abstract
LLM agents interact with external systems by generating structured tool calls. Given a user request, conversational context, and a catalog of tool schemas, a tool-calling model must select tools and generate their arguments, potentially producing multiple calls in a single response. Standard autoregressive decoding generates these calls token by token, incurring substantial latency for requests involving multiple calls or many argument fields. The explicit argument structure offers opportunities for parallel generation, but later argument values may depend on preceding fields and calls, so independently generated values can differ from the target model's output. We present SchemaFill, a framework for efficient LLM tool calling through slot-parallel speculative decoding. SchemaFill generates future slot values concurrently as candidates, without requiring advance knowledge of the actual call sequence or argument values. Candidates spanning multiple fields and calls are concatenated for verification by the target model under the actual output prefix. Only verified tokens are committed, and the target supplies corrections when candidates disagree. This applies target verification while exploiting parallelism across slots and calls. On Glaive and BFCL, SchemaFill achieves up to a 4.05$\times$ improvement in end-to-end throughput over autoregressive decoding. Code is available at https://github.com/Czzzk/SchemaFill.
Problem

Research questions and friction points this paper is trying to address.

Tool Calling
Autoregressive Decoding
Latency
Structured Generation
Parallel Generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speculative Decoding
Tool Calling
Slot-Parallel Generation
LLM Agents
Throughput Optimization
Z
Zhi-Kai Chen
School of Artificial Intelligence, Nanjing University, China; National Key Laboratory for Novel Software Technology, Nanjing University, China
S
Song-Yan Li
Nanjing University, China
De-Chuan Zhan
De-Chuan Zhan
Nanjing University, China
Machine LearningData Mining
Han-Jia Ye
Han-Jia Ye
Nanjing University
Machine LearningData MiningMetric LearningMeta-Learning