π€ AI Summary
This study addresses the prohibitive computational overhead of high-resolution video generation and the subsampling bottleneck inherent in conventional multi-step refinement approaches by proposing a one-step video refinement framework. The method introduces a three-stage training paradigm: high-resolution continual pre-training to enhance model capacity, reinforcement learning-based post-training to optimize generation quality, and one-step distillation to overcome inference speed limitations. Additionally, a dedicated evaluation benchmark, Refiner-Bench, is constructed. Experimental results demonstrate that the proposed approach achieves high-quality, single-step upscaling from low-resolution inputs to 4K videos. It outperforms existing external refiners on metrics such as VBench while accelerating inference latency by 8.91Γ, thereby significantly improving both the efficiency and fidelity of 4K video generation.
π Abstract
High-resolution video generation is expensive, as its cost grows rapidly with the number of spatiotemporal tokens. A practical alternative first generates a lower-resolution video and then applies a refiner, but conventional multi-step refinement introduces a second sampling bottleneck. We present SoL-Refiner, a one-step video refiner that transforms low-resolution model outputs into 4K videos with a single denoising step. Our three-stage recipe combines high-resolution continual training, reinforcement learning (RL) post-training, and a final one-step distillation. We introduce Refiner-Bench, a video refinement benchmark constructed from the outputs of different video generators, and use a shared-input protocol to compare refiners at approximately 2K output resolution. At 2K, the one-step SoL-Refiner outperforms all external refiners on the VBench and UniPercept averages, while at $3840\!\times\!2176$ it improves both metrics over the three-step LTX-2.3 Refiner. With the complete acceleration stack, SoL-Refiner achieves an $8.91\times$ speedup in refinement latency over the same baseline in our 2K latency setting.