OmniTide: Co-Designing Algorithms and Systems for Efficient On-Device Omni-LLM Streaming

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the memory and computational bottlenecks caused by continuously growing KV caches during on-device, all-modality streaming inference. To this end, it introduces a pioneering algorithm-system co-design framework. Methodologically, the work proposes a unit-based abstraction mechanism alongside an OmniPick logic retention strategy to preserve critical multimodal context. Furthermore, memory management is optimized through OmniPage physical partitioning, dynamic sparse token compression, and modality-aware structured sparsity. Experimental results demonstrate that the proposed system achieves a 12.72× kernel speedup, a 2.4× reduction in latency, an 18% improvement in accuracy, and a 26.7% decrease in physical KV span. Ultimately, this work establishes a novel paradigm for efficient, real-time multimodal inference on edge devices.
📝 Abstract
On-device streaming omni-modal inference safeguards user privacy and eliminates prohibitive per-token API costs, but faces a critical bottleneck: the continuous influx of multimodal data rapidly exhausts constrained memory and compute budgets via monotonic KV cache growth. Existing sparse attention methods fall short, either incurring prohibitive online estimation latency or destroying interleaved cross-modal context, while failing to resolve physical memory fragmentation. We present OmniTide, the first algorithm-system co-design tailored for efficient on-device streaming omni-modal inference. Driven by the observation of modality-aware structural sparsity, OmniTide adopts a unit-based abstraction with two components: (1) At the algorithm level, OmniPick logically retains critical multimodal context based on unit boundaries and modality importance to preserve task accuracy; (2) At the system level, OmniPage physically partitions the cache by retention likelihood and dynamically compacts surviving sparse tokens, minimizing both memory fragmentation and data-movement overhead. Extensive evaluations across three streaming benchmarks and two consumer-device architectures show that OmniTide achieves up to $12.72\times$ kernel speedups and $2.40\times$ lower stream-loop latency. On StreamingBench, it improves accuracy by up to 18.0 percentage points over sliding-window baselines at comparable session cost. OmniPage further reduces the physical KV span by up to 26.7% relative to native logical eviction, unlocking real-time, infinite-context streaming on edge devices.
Problem

Research questions and friction points this paper is trying to address.

On-device inference
Omni-modal streaming
KV cache
Memory fragmentation
Sparse attention
Innovation

Methods, ideas, or system contributions that make the work stand out.

algorithm-system co-design
omni-modal streaming inference
sparse attention
KV cache management
on-device LLM