DuplexCadence: Exact State and Execution from a Speech Model's Declared Timelines

πŸ“… 2026-09-28
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Full-duplex speech models suffer from response timeouts and memory inefficiency due to the tail overhead inherent in audio synthesis. This work proposes a runtime optimization mechanism grounded in the model’s native clock, which achieves on-demand allocation and precise replay through GPU orchestration, static memory analysis, and computation graph recording, thereby eliminating scheduling gaps. Furthermore, streaming state management is dynamically optimized to reduce memory footprint. Experimental results demonstrate that the proposed approach accelerates inference by 2.85Γ— and reduces peak memory consumption by 38.8%, ensuring stable operation within a one-second interaction cycle.
πŸ“ Abstract
Full-duplex speech models support streaming interaction that listens and speaks at the same time. Serving them is governed by a strict, repeating deadline: conversation advances on a one-second cadence, and every second of input must be turned into a second of speech before the next second arrives. Because stages within a session run in strict sequence, per-invocation overhead cannot be batched away. Profiling reveals that the autoregressive stages of a duplex second already fit within the period, whereas the token-to-audio synthesis tail is what causes overruns. This tail stage suffers from orchestration slack where the GPU is left waiting as thousands of tiny, regular operations are issued one by one, while also wasting substantial memory by over-provisioning state at static implementation constants. Existing remedies, such as graph recording and demand-sized allocation, fail because streaming state dynamics violate their prerequisites. The root cause is that the runtime lacks the model's native clocks: the per-region counters that govern advancement rates and retention policies. We propose DuplexCadence, which explicitly declares native clocks to the runtime and derives two mutually enabling rules: demand-sized state allocation at a stable address, and exact-shape graph replay without padding. The former eliminates idle memory and stabilizes tensor pointers, while the latter removes orchestration slack without padding overhead. Evaluated on four released models across three decoder architectures with bit-for-bit identical output, DuplexCadence reaches $2.85\times$ the stock runtime's speed at $38.8\%$ lower peak memory. On the live duplex path, mean SPEAK time falls from $14\%$ over the one-second cadence to $2\%$ under it, enabling models to reliably keep up with interactive speech while markedly expanding multi-
Problem

Research questions and friction points this paper is trying to address.

full-duplex speech models
real-time streaming inference
token-to-audio synthesis latency
orchestration slack
memory over-provisioning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Full-duplex speech models
Native clocks declaration
Demand-sized state allocation
Exact-shape graph replay
Streaming inference optimization
πŸ”Ž Similar Papers
No similar papers found.