Group-of-Latents: Perceptual Video Compression at Extreme Bitrates via Masked Latent Generative Modeling

πŸ“… 2026-07-21
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the challenge of achieving both high perceptual quality and temporal consistency in video compression at extremely low bitrates (<0.005 bpp), where existing methods struggle. The authors propose a unified generative framework that introduces a causal tokenizer to decompose latent representations into I-latents and P-latents, and employs a Group-of-Latents strategy for structured modeling. Key latents are efficiently encoded via an I-frame Deep Compression Module (I-DCM), while a unified latent denoising module (U-LDM), built upon a pre-trained Diffusion Transformer (DiT), reconstructs high-fidelity intra-frame textures and coherent temporal dynamics directly from noise. Notably, this approach incurs no additional bitrate overhead and significantly outperforms current state-of-the-art methods under extreme low-bitrate constraints, delivering spatially detailed and temporally stable visual quality.
πŸ“ Abstract
Most existing video compression algorithms follow a paradigm of transformation and quantization, optimizing the trade-off between distortion and bitrate. However, extremely low-bitrate compression remains an underexplored frontier where perceptual quality optimization under severely constrained coding resources has not been adequately addressed. In this paper, we propose a unified generative framework that leverages pre-trained Diffusion Transformer (DiT) priors to achieve high perceptual quality at extremely low bitrates. We first introduce a flexible Group-of-Latents (GoL) strategy within the latent space of a causal tokenizer, explicitly partitioning the latent stream into intra $I$-latents and inter $P$-latents. The Deep Compression Module (I-DCM) then encodes key $I$-latents to preserve perceptual anchors with minimal overhead. Building upon these anchors, the DiT-based Unified Latent Denoising Module (U-LDM) refines intra-frame textures and synthesizes $P$-latents from noise, reconstructing temporal dynamics at zero additional bitrate cost. Extensive experiments demonstrate that our method uniquely operates in the extreme-low-bitrate regime (e.g., (<0.005) bpp), achieving state-of-the-art perceptual fidelity with rich spatial details and robust temporal consistency. The code will be made publicly available.
Problem

Research questions and friction points this paper is trying to address.

extreme low-bitrate
perceptual video compression
latent generative modeling
temporal consistency
spatial details
Innovation

Methods, ideas, or system contributions that make the work stand out.

Group-of-Latents
Diffusion Transformer
Perceptual Video Compression
Extreme Low Bitrate
Latent Generative Modeling
Shaokang Wang
Shaokang Wang
Infinera; UMBC
Mathematical ModelingComputational PhysicsStatistical Signal Processing
J
Jinchang Xu
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University, Beijing, China
P
Peidong Jia
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University, Beijing, China
Z
Zhijian Hao
State Key Discipline Laboratory of Wide Band-Gap Semiconductor Technology, School of Microelectronics, Xidian University, Xi’an, China
S
Siyuan Qian
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University, Beijing, China
Fei Zhao
Fei Zhao
School of Computer Science, Inner Mongolia University
Acoustic Echo CancellationActive Noise Control
Rui Ma
Rui Ma
Peking University
AIGCCopyright ProtectionDiffraction Imaging
X
Xiaozhu Ju
X-humanoid, Beijing, China
J
Jian Tang
X-humanoid, Beijing, China
Xiaodong Xie
Xiaodong Xie
Peking University
media processorVLSI
Shanghang Zhang
Shanghang Zhang
Peking University
Embodied AIFoundation Models
H
Huizhu Jia
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University, Beijing, China