🤖 AI Summary
Full-duplex speech models suffer from excessive memory overhead during prolonged interactions due to the continuous accumulation of acoustic key-value (KV) states. This work proposes an acoustic-to-text KV compression mechanism that converts speech into textual memory during listening pauses and evicts obsolete acoustic states when the memory budget is exceeded. Knowledge distillation is incorporated to preserve native listening and reading behaviors, while LoRA fine-tuning and streaming cache management are employed to optimize training. The proposed approach reduces peak KV cache usage by 64.6% and significantly improves transcription and question-answering performance, while maintaining conversational interaction capabilities on par with the baseline.
📝 Abstract
Full-duplex speech language models continuously accumulate acoustic key-value (KV) states, making long-running interactions memory-intensive. During listening, the model can finish processing an audio unit before the next arrives; we term the remaining interval listening-time slack. We propose acoustic-to-text KV compression, which introduces a transcription side channel to convert incoming speech into compact textual memory within this interval. When the cache exceeds a target budget during inference, older acoustic states are evicted while transcripts and recent acoustic context remain. We train the side channel with LoRA using cross-entropy on transcription segments. To preserve listening and speaking behavior, we apply knowledge distillation to the original model's token-level output distributions at native prediction positions. On ten-minute LongSpeech sessions, our MiniCPM-o 4.5 implementation reduces peak streaming KV-cache size by 64.6% compared with the same model without eviction. The proposed method also improves transcription, temporal question answering, and summarization over the baseline. Full-Duplex-Bench evaluations further show comparable pause-handling, turn-taking, and interruption performance.