🤖 AI Summary
This work addresses the parameter redundancy and limited generalization gains of existing unified ultra-high-definition (UHD) video restoration models by proposing a recurrent mixture-of-experts architecture. Specifically, we design a degradation-aware low-rank embedding module and an iterative spatiotemporal adaptive normalization module, while introducing an input-conditioned dynamic depth prediction mechanism. These components collectively establish a lightweight recurrent learning paradigm capable of efficiently handling cross-domain restoration tasks. Requiring only approximately 0.884M trainable parameters, the proposed model achieves state-of-the-art performance on both benchmark datasets and real-world scenarios for tasks such as dehazing and deraining. Ultimately, this approach enables the unified and efficient processing of diverse UHD video restoration tasks with minimal computational overhead.
📝 Abstract
Recently, unified high-definition image restoration has attracted considerable attention; however, existing models tend to excessively increase their depth in pursuit of improved generalization, which often yields only limited gains. Meanwhile, loop-based learning paradigms have drawn widespread attention due to their low parameter counts and strong regression capability, as exemplified by GPT-6 and looped Transformers. In this paper, we introduce the loop learning paradigm to address restoration tasks that require cross-domain learning. Specifically, we propose LoopMoEVR, a loop-based mixture-of-experts model capable of handling degraded ultra-high-definition (UHD) inputs. First, a degradation-conditioned low-rank loop embedding is designed to construct input-dependent stage conditions. Second, a spatio-temporal iterative adaptive normalization module, termed IterAda3DN, is developed to fuse local features with global loop context, thereby performing position-wise affine modulation. Finally, the expert branches further integrate the attention-updated local and global video states with the loop conditions to generate dedicated modulation parameters, while an input-conditioned depth predictor adaptively configures the number of loop iterations. With only approximately 0.884M trainable parameters, the proposed model uniformly handles UHD video dehazing, deraining, denoising, and low-light enhancement tasks, achieving state-of-the-art restoration performance on both public benchmarks and real-world scenarios.