๐ค AI Summary
This study addresses the high latency of iterative sampling and insufficient temporal consistency in diffusion-based real-time video super-resolution by proposing a streaming inference framework built upon latent diffusion models. The core innovations include a cross-step attention mechanism that reuses intermediate denoising features to circumvent explicit temporal modeling, and a trajectory-coupled diffusion scheduling strategy that reduces computational complexity from O(NยทS) to O(N+S). Experimental results demonstrate that, following an initial cold start, the proposed method achieves frame rates exceeding 40 FPS at 512ร512 resolution while substantially enhancing both perceptual quality and temporal realism, thereby enabling efficient and high-fidelity video super-resolution inference.
๐ Abstract
Real-time video super-resolution requires high spatio-temporal fidelity under strict latency constraints, challenging diffusion models due to their iterative sampling cost and limited temporal coordination. We propose a streaming-aware framework that adapts pretrained single-image latent diffusion models for efficient video super-resolution (VSR) by exploiting the sequential structure of video streams. Our Cross-Step Attention mechanism reuses intermediate denoising features across adjacent frames and diffusion steps, enabling temporal information exchange without explicit temporal modeling. We further introduce Trajectory-Coupled Diffusion Scheduling, which aligns adjacent diffusion states and provides cleaner intermediate representations for cross-step conditioning, improving temporal coherence. These components are integrated into a streaming inference pipeline that incrementally propagates latent states across frames, reducing the effective computational complexity from $O(N \cdot S)$ to $O(N + S)$ for $N$ frames and $S$ diffusion steps. Experiments on REDS4 and YouHQ40-Test demonstrate improved perceptual quality and temporal realism while maintaining frame-wise stability. Our method achieves over 40 FPS at $512 \times 512$ resolution after cold start, enabling real-time VSR without explicit temporal modeling.