A Spatiotemporal Semantic Importance-Guided Unified Compression and Editing Framework for AI-Generated Videos

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the substantial storage and transmission overhead of AI-generated videos and the challenge of preserving semantic consistency during editing. We propose a unified compression and editing framework built upon frozen generative priors. The core innovation lies in a spatiotemporal semantic importance guidance mechanism that distinguishes semantic invariants from regenerable details. By integrating latent space projection, adaptive non-uniform bit allocation, and continuous side information refinement, our approach achieves efficient coding while effectively reusing generative priors. This method significantly reduces data volume while enabling high-fidelity video reconstruction and structure-consistent, prompt-driven editing, thereby unifying the compression and editing pipelines into a single coherent framework.
📝 Abstract
AI-generated videos are rapidly increasing in volume, duration, and resolution, creating growing demands for efficient storage and transmission. Unlike natural videos captured from the physical world, AI-generated videos are samples from a learned generative distribution, where semantic structures are critical to content consistency, while many local textures and stochastic details can be plausibly regenerated. This distinction suggests that compression should preserve semantically important spatiotemporal information rather than reconstruct every pixel of a particular generative sample. Beyond reconstruction, AI-generated videos also create a practical need for prompt-based editing, where users expect to modify generated content while preserving its original spatiotemporal semantics. Motivated by these observations, we propose a unified compression and editing framework for AI-generated videos that incorporates a frozen video generator as a reusable generative prior. Within this framework, we design three spatiotemporal semantic importance-guided techniques that respectively address what to transmit, how much to transmit, and how to use the transmitted side information. First, an innovation selection method projects the latent discrepancy using spatiotemporal semantic importance, so that the selected innovations prioritize semantic invariants over replaceable generative variations. Second, a frame-adaptive bit allocation method estimates the nonuniform semantic demands of latent frames and allocates more innovations to frames requiring stronger semantic preservation. Third, a unified reconstruction and editing method continuously adjusts the influence of the transmitted side information, enabling the same compressed representation to provide strong guidance for faithful reconstruction or serve as a flexible semantic anchor for structure-preserving prompt-driven editing.
Problem

Research questions and friction points this paper is trying to address.

AI-generated video compression
video storage and transmission
prompt-based video editing
spatiotemporal semantic preservation
Innovation

Methods, ideas, or system contributions that make the work stand out.

AI-generated video compression
spatiotemporal semantic importance
generative prior
unified compression and editing
frame-adaptive bit allocation