🤖 AI Summary
This study addresses the lack of workflow awareness in LLM serving caused by information fragmentation between agent frameworks and inference engines. To this end, it proposes HEAR, a bidirectional protocol that establishes, for the first time, semantic separation and two-way communication between orchestrators and engines. By standardizing the linkage between workflow intent and runtime state, HEAR enables cache-aware coordination and load-adaptive execution mode selection, supporting diverse scheduling strategies without modifying model or task definitions. Experimental results demonstrate that under memory-constrained concurrent scenarios, the proposed approach achieves a 1.61× improvement in batch throughput, a 2.23× reduction in time-to-first-token latency, and a 2.45× end-to-end speedup, all without compromising generation quality.
📝 Abstract
LLM agents increasingly execute complex workflows involving multi-turn reasoning, tool use, and parallel agents. Efficient serving requires decisions that span two layers with complementary information: the agent harness understands workflow dependencies, context lifecycles, and execution objectives, whereas the inference engine observes request queues, KV-cache state, resource pressure, and execution capabilities. Existing interfaces do not systematically connect these views, limiting workflow-aware execution. HEAR, a bidirectional Harness--Engine Pairing protocol for agentic LLM serving. HEAR standardizes how the harness communicates workflow intent and execution requirements and how the engine returns runtime state, capabilities, and outcomes. By separating protocol semantics from optimization policies, HEAR supports diverse coordination strategies without changing workflow or model semantics. We instantiate HEAR for online cache-aware runtime coordination and workload-aware execution-mode selection for agent roles. Across four conversational and research-agent benchmarks under memory-constrained, concurrent serving, HEAR achieves a $1.61\times$ batch speedup and reduces median time-to-first-token by $2.23\times$ on SCBench. Mooncake shows that workflow intent and live engine state provide complementary benefits across load regimes. On BrowseComp-Plus and DeepResearchBench, workload-specific configurations yield $1.23\times$ and $2.45\times$ end-to-end speedups, respectively, without observed task-quality degradation. These results establish HEAR as a reusable coordination substrate for efficient agentic LLM serving.