Compressing Scene Dynamics: A Generative Approach

📅 2024-10-13
🏛️ Data Compression Conference
📈 Citations: 3
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses ultra-low-bitrate compression of dynamically varying scene videos. Instead of conventional content-based modeling, it leverages natural motion patterns—such as flower swaying or boat drifting—as priors. Methodologically, it introduces, for the first time, a lightweight generative prior for scene motion, establishing a novel framework comprising dense motion representation, sparse motion coding, and optical-flow-guided diffusion decoding—fully abandoning inter-frame prediction. Key contributions are: (1) learning compact, generalizable motion priors from common dynamic scenes; and (2) designing a flow-driven generative decoding mechanism enabling high-fidelity dynamic reconstruction. Experiments demonstrate substantial gains over VVC across diverse dynamic sequences, maintaining strong motion consistency and visual quality at ultra-low bitrates (0.01–0.1 bpp), with comprehensive improvements in rate-distortion performance.

Technology Category

Computer Vision: Motion & TrackingMachine Learning: Deep Generative Models & AutoencodersSearch and Optimization: Sampling/Simulation-based Search

Application Category

Graph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphsUser Modeling, Personalization and Recommendation: On-Device user modeling, personalization, and recommendationSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
📝 Abstract
This paper proposes to learn generative priors from the motion patterns instead of video contents for generative video compression. The priors are derived from small motion dynamics in common scenes such as swinging trees in the wind and floating boat on the sea. Utilizing such compact motion priors, a novel generative scene dynamics compression framework is built to realize ultra-low bit-rate communication and high-quality reconstruction for diverse scene contents. At the encoder side, motion priors are characterized into compact representations in a dense-to-sparse manner. At the decoder side, the decoded motion priors serve as the trajectory hints for scene dynamics reconstruction via a diffusion based flow-driven generator. The experimental results illustrate that the proposed method can achieve superior rate-distortion performance and outperform the state-of-the-art conventional video codec Versatile Video Coding (VVC) on scene dynamics sequences.
Problem

Research questions and friction points this paper is trying to address.

Compresses scene dynamics using motion pattern priors
Enables ultra-low bitrate video communication
Achieves superior rate-distortion performance over ECM
Innovation

Methods, ideas, or system contributions that make the work stand out.

Leverages motion pattern priors for compression
Uses dense-to-sparse transformation at encoder
Employs flow-driven diffusion model at decoder
City University of Hong Kong | DAMO Academy, Alibaba Group
S
Shanzhi Yin
City University of Hong Kong
Z
Zihan Zhang
City University of Hong Kong
B
Bo Chen
City University of Hong Kong
S
Shiqi Wang
City University of Hong Kong
Yan Ye
Yan Ye
Alibaba Inc
video coding