Free-GVC: Towards Training-Free Extreme Generative Video Compression with Temporal Coherence

📅 2026-02-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of modeling temporal correlations in generative video compression at extremely low bitrates, where existing methods often suffer from severe flickering and temporal inconsistency. We propose the first training-free generative video compression framework, reframing reconstruction as a latent trajectory optimization problem guided by a video diffusion prior. Our approach introduces adaptive quality control and cross-GOP latent alignment at the group-of-pictures (GOP) level, coupled with an online rate-aware proxy model and cross-GOP latent fusion to significantly enhance temporal coherence and visual fidelity. Experimental results demonstrate that the proposed method achieves a 93.29% average BD-Rate reduction over DCVC-RT in terms of the DISTS metric, and user studies confirm its superior perceptual performance under ultra-low-bitrate conditions.

Technology Category

Computer Vision: Diffusion Models for VisionMachine Learning: Learning on the Edge & Model CompressionNatural Language Processing: Generation

Application Category

Economics, Online Markets and Human Computation: Economic ramifications for generative AI infrastructure and applicationsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and ranking
📝 Abstract
Building on recent advances in video generation, generative video compression has emerged as a new paradigm for achieving visually pleasing reconstructions. However, existing methods exhibit limited exploitation of temporal correlations, causing noticeable flicker and degraded temporal coherence at ultra-low bitrates. In this paper, we propose Free-GVC, a training-free generative video compression framework that reformulates video coding as latent trajectory compression guided by a video diffusion prior. Our method operates at the group-of-pictures (GOP) level, encoding video segments into a compact latent space and progressively compressing them along the diffusion trajectory. To ensure perceptually consistent reconstruction across GOPs, we introduce an Adaptive Quality Control module that dynamically constructs an online rate-perception surrogate model to predict the optimal diffusion step for each GOP. In addition, an Inter-GOP Alignment module establishes frame overlap and performs latent fusion between adjacent groups, thereby mitigating flicker and enhancing temporal coherence. Experiments show that Free-GVC achieves an average of 93.29% BD-Rate reduction in DISTS over the latest neural codec DCVC-RT, and a user study further confirms its superior perceptual quality and temporal coherence at ultra-low bitrates.
Problem

Research questions and friction points this paper is trying to address.

generative video compression
temporal coherence
flicker
ultra-low bitrates
video diffusion
Innovation

Methods, ideas, or system contributions that make the work stand out.

training-free
generative video compression
temporal coherence
diffusion prior
latent trajectory compression
🔎 Similar Papers
No similar papers found.