π€ AI Summary
Video compression reduces transmission overhead but degrades visual quality, leading to significant performance drops in downstream vision tasks (e.g., object detection, action recognition). To address this, we propose a codec-aware, plug-and-play video enhancement framework that requires no modification to existing codecs and introduces zero inference latency. Our approach features two key innovations: (1) a novel hierarchical compressed-domain awareness mechanism that jointly models spatiotemporal priors from bitstreams; and (2) co-optimized lightweight frame-level enhancement (BAE) and compression-aware adaptation (CAA) networks, enabling zero-intrusion, cross-standard adaptability across H.264, H.265, and AV1. Evaluated on multiple benchmarks, our method consistently outperforms state-of-the-art approaches, improving average downstream task accuracy by 3.2β5.7% while maintaining real-time deployment capability.
π Abstract
As a widely adopted technique in data transmission, video compression effectively reduces the size of files, making it possible for real-time cloud computing. However, it comes at the cost of visual quality, posing challenges to the robustness of downstream vision models. In this work, we present a versatile codec-aware enhancement framework that reuses codec information to adaptively enhance videos under different compression settings, assisting various downstream vision tasks without introducing computation bottleneck. Specifically, the proposed codec-aware framework consists of a compression-aware adaptation (CAA) network that employs a hierarchical adaptation mechanism to estimate parameters of the frame-wise enhancement network, namely the bitstream-aware enhancement (BAE) network. The BAE network further leverages temporal and spatial priors embedded in the bitstream to effectively improve the quality of compressed input frames. Extensive experimental results demonstrate the superior quality enhancement performance of our framework over existing enhancement methods, as well as its versatility in assisting multiple downstream tasks on compressed videos as a plug-and-play module. Code and models are available at https://huimin-zeng.github.io/PnP-VCVE/.