AiFlow: Token-Native Reactive Orchestration with Bounded Backpressure for Streaming LLM Applications

📅 2026-08-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing LLM streaming frameworks lack unified support for token-level flow control, queue management, scheduling, and backpressure, often relying on ad hoc callbacks that lead to system complexity and unreliability. This work proposes AiFlow—the first token-native reactive orchestration model—which formalizes LLM generation streams as typed Context<T> events propagated through a directed stream graph. Node guardians uniformly manage queue boundaries, concurrency, ordering, and fault tolerance. AiFlow supports DSL/JSON graph compilation and provides static verification for type safety, stateful concurrency, and injection compatibility. Its reactive architecture, combined with runtime-enforced policies, guarantees bounded memory usage. Experiments demonstrate a 70.9–94.7% reduction in time-to-first-token latency, stable and controllable queue depths, and a 93.7–96.5% decrease in maximum queue length.
📝 Abstract
Large language model (LLM) applications increasingly operate as streaming workflows combining retrieval, tool calls, safety filters, and multi-agent coordination. Although contemporary frameworks expose provider deltas, workflow nodes often treat generation as coarse request-response steps, leaving queue management, worker allocation, ordering, and backpressure to ad hoc callback code. This paper presents AiFlow, a token-native reactive orchestration model that normalizes provider deltas into typed Context<T> events propagated through a directed streaming graph. Each node is managed by a Node Guardian that declares and enforces local queue bounds, worker concurrency, ordering, overflow policy, cancellation propagation, and retry discipline. We formalize the bounded-memory property, present the compilation from a compact DSL and JSON graph form, and provide static validation for type safety, state concurrency, and injection compatibility. Controlled microbenchmarks, captured DeepSeek trace replay (30 runs), descriptive online runs, LangGraph baselines, a streaming RAG workload, and an Ollama local-backend check show that AiFlow does not alter provider-side Model TTFT but reduces Application TTFPT by 70.9-94.7\% versus aggregation and keeps runtime-owned queue depth within declared bounds (93.7-96.5\% MaxQ reduction versus unbounded policies). The supplementary artifact contains scripts, raw traces, machine-readable tables, checksums, and an API-free smoke test; the public implementation is available through the FIT Framework repository.
Problem

Research questions and friction points this paper is trying to address.

streaming LLM applications
backpressure
reactive orchestration
queue management
workflow coordination
Innovation

Methods, ideas, or system contributions that make the work stand out.

token-native
reactive orchestration
bounded backpressure
streaming LLM
Node Guardian
🔎 Similar Papers