RelayVSR: Large-Small Model Collaboration for Efficient Real-World Video Super-Resolution

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inherent trade-off between the prohibitive computational cost of large-model video super-resolution and the insufficient detail reconstruction of small models by proposing a streaming collaborative framework. Methodologically, this work introduces a Dual-Memory Video Transformer and Video-Aware Reference Optimization (VARO) to synergize large and small models via a sparse generative relay mechanism, while employing a dual-level reinforcement learning reward scheme to mitigate error accumulation. Experimental results demonstrate that the proposed framework achieves 29.29 FPS at 1080p resolution with only 13.82 GB of VRAM consumption and a first-frame latency of 0.327 seconds. These metrics significantly outperform existing methods, establishing a novel paradigm for efficient real-world video super-resolution.
📝 Abstract
Large generative models can recover realistic detail in real-world video super-resolution (VSR), but processing an entire video with them is computationally expensive. In this work, we present RelayVSR, a streaming VSR framework built on the Sparse Generative Relay mechanism. A large generative model generates reference latents for sparse keyframes, while a lightweight VSR network uses these references and low-resolution video to super-resolve every frame. The lightweight VSR network, implemented as a Dual-Memory Video Transformer, reuses keyframe information across frames and updates recent video context, supporting first-keyframe conditioning and dual-endpoint conditioning with bounded lookahead. However, errors in shared keyframes can propagate and accumulate across output frames, making keyframe quality alone an insufficient optimization target. We address this collaboration gap with Video-Aware Reference Optimization (VARO), which uses reinforcement learning to update the large generative model with two reward levels: a system-level reward evaluates videos produced by the fixed lightweight VSR network, while a reference-level reward evaluates decoded keyframe quality. VARO improves final video quality over direct joint training, and its dual-level rewards outperform a system-level reward alone. At 1080p on a single NVIDIA A100 80GB, dual-endpoint RelayVSR with a 15-frame keyframe interval reaches 29.29 FPS, 13.82 GB peak GPU memory, and 0.327 s first-frame model latency, compared with 7.80 FPS, 24.447 GB, and 2.83 s for FlashVSR-Tiny. The code is available at https://github.com/kopperx/RelayVSR.
Problem

Research questions and friction points this paper is trying to address.

Video Super-Resolution
Computational Efficiency
Error Propagation
Large-Small Model Collaboration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Video Super-Resolution
Sparse Generative Relay
Dual-Memory Video Transformer
Reinforcement Learning
Large-Small Model Collaboration
🔎 Similar Papers
No similar papers found.