🤖 AI Summary
This work addresses the challenge of achieving visually usable, efficient, and robust video communication under ultra-low bandwidth and poor network conditions, where conventional approaches struggle to balance perceptual quality, transmission efficiency, and resilience. The authors propose a generative video communication framework that departs from pixel-level fidelity paradigms by leveraging task-oriented perceptual reconstruction and generative priors. This framework jointly optimizes bandwidth, computational, and memory resources through cross-segment memory reuse and runtime state sharing mechanisms. These innovations substantially reduce transmission overhead while simultaneously enhancing decoding speed, link robustness, and perceptual quality at extremely low bitrates, thereby enabling efficient and practically viable video transmission in constrained environments.
📝 Abstract
Under the AI Flow framework, communication is shifting from transmitting fidelity-oriented information flows toward delivering task-oriented and perception-oriented token flows across heterogeneous network resources. Video communication is a fundamental component of modern information networks. However, under ultra-low-bandwidth and weak-network conditions, conventional video coding and transmission methods, which are primarily optimized for pixel-level fidelity, often struggle to balance visual usability, transmission efficiency, and robustness to unstable links. With the rapid advancement of generativemodels, video communication is also moving from precise signal reconstruction toward receiver-side perceptual utility and system-level usability. In this paper, we propose Generative Transmission (GenTrans) for video communication under ultra-low-bandwidth and weak-network conditions. Built upon Generative Video Compression (GVC), GenTrans formulates video transmission as a joint optimization problem involving bandwidth, computation, and memory, rather than treating it merely as a signal coding task. By leveraging generative priors, cross-clip memory reuse, runtime state reuse, and weak-network-aware transport, GenTrans significantly reduces transmission overhead while enabling visually coherent and practically useful reconstruction. Experimental results show that GenTrans supports effective video transmission under ultra-low-bitrate and weak-network conditions, achieving improved transmission efficiency, decoding efficiency, and robustness while preserving perceptual quality.