DynBranch: Speculative Subgraph Reuse for Dynamic Agentic LLM Serving

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the computational serialization bottleneck, termed the branch resolution barrier, caused by unresolved branches in LLM agent workflows. We propose a speculative execution framework that requires no modifications to existing engines. The core methodology introduces a stable coordinate mechanism that renders unresolved branches addressable, thereby enabling the preemptive execution and result reuse of candidate subgraphs. Furthermore, a two-level adaptive controller is constructed to dynamically schedule speculative tasks based on expected utility and overhead costs. Experiments conducted with Qwen3-series models across multiple hardware platforms demonstrate that the proposed approach reduces average latency by 32% compared to the strongest baseline and by 46%–66% relative to a non-reuse benchmark, while strictly preserving workflow semantic correctness.
📝 Abstract
Agentic LLM workflows decide their execution paths at runtime. Downstream computation may be predictable, or may have run before, yet it cannot begin until the model or the user resolves the branch. We call this serialization the branch-resolution barrier. Caching alone does not hide it: the key that identifies a reusable result is not known until then. In this paper, we propose DynBranch, which makes an unresolved branch addressable before it resolves. Its stable coordinate lets candidate subgraphs run during resolution and completed subgraph results be reused across later requests. A two-level controller admits this work when its expected benefit exceeds the load price. DynBranch sits at the model-API boundary and requires no changes to agent harnesses or model execution engines. Across four agentic workloads with Qwen3-32B on 4x H200 GPUs, DynBranch reduces mean latency by up to 32% over each workload's strongest prior system and by 46-66% against a no-reuse floor, while preserving workflow results. The benefit persists across backbone families and on a commodity Qwen3-8B/RTX 4090 deployment.
Problem

Research questions and friction points this paper is trying to address.

Agentic LLM
Branch-resolution barrier
Dynamic workflow serving
Subgraph reuse
Latency optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speculative Subgraph Reuse
Dynamic Agentic LLM Serving
Branch-Resolution Barrier
Stable Coordinate
Two-Level Controller
🔎 Similar Papers
No similar papers found.