🤖 AI Summary
This work addresses ultra-low-bitrate compression of dynamically varying scene videos. Instead of conventional content-based modeling, it leverages natural motion patterns—such as flower swaying or boat drifting—as priors. Methodologically, it introduces, for the first time, a lightweight generative prior for scene motion, establishing a novel framework comprising dense motion representation, sparse motion coding, and optical-flow-guided diffusion decoding—fully abandoning inter-frame prediction. Key contributions are: (1) learning compact, generalizable motion priors from common dynamic scenes; and (2) designing a flow-driven generative decoding mechanism enabling high-fidelity dynamic reconstruction. Experiments demonstrate substantial gains over VVC across diverse dynamic sequences, maintaining strong motion consistency and visual quality at ultra-low bitrates (0.01–0.1 bpp), with comprehensive improvements in rate-distortion performance.
📝 Abstract
This paper proposes to learn generative priors from the motion patterns instead of video contents for generative video compression. The priors are derived from small motion dynamics in common scenes such as swinging trees in the wind and floating boat on the sea. Utilizing such compact motion priors, a novel generative scene dynamics compression framework is built to realize ultra-low bit-rate communication and high-quality reconstruction for diverse scene contents. At the encoder side, motion priors are characterized into compact representations in a dense-to-sparse manner. At the decoder side, the decoded motion priors serve as the trajectory hints for scene dynamics reconstruction via a diffusion based flow-driven generator. The experimental results illustrate that the proposed method can achieve superior rate-distortion performance and outperform the state-of-the-art conventional video codec Versatile Video Coding (VVC) on scene dynamics sequences.