🤖 AI Summary
Existing image/video compression methods prioritize human visual perception, often compromising downstream machine vision performance. To address this, we propose the Efficient Adaptive Compression (EAC) framework—the first to jointly optimize for both human visual fidelity and accuracy in downstream vision tasks (e.g., object detection and semantic segmentation). Methodologically, EAC introduces a dual-module co-design: (1) adaptive latent feature subset selection to dynamically balance multi-objective optimization, and (2) a lightweight task adapter leveraging delta-tuning to unlock the potential of pre-trained vision models without full fine-tuning. Technically, EAC integrates foundational neural image coding (NIC) architectures (Ballé et al., 2018; Cheng et al., 2020) and neural video coding (NVC) frameworks (DVC/FVC), enhanced with latent-space adaptive pruning and parameter-efficient adaptation. Evaluated on VOC2007, COCO, and DAVIS benchmarks, EAC significantly improves mAP and mIoU for machine vision tasks while preserving PSNR and MS-SSIM—human-perceived quality metrics—with only marginal bitrate overhead.
📝 Abstract
While most existing neural image compression (NIC) and neural video compression (NVC) methodologies have achieved remarkable success, their optimization is primarily focused on human visual perception. However, with the rapid development of artificial intelligence, many images and videos will be used for various machine vision tasks. Consequently, such existing compression methodologies cannot achieve competitive performance in machine vision. In this work, we introduce an efficient adaptive compression (EAC) method tailored for both human perception and multiple machine vision tasks. Our method involves two key modules: 1), an adaptive compression mechanism, that adaptively selects several subsets from latent features to balance the optimizations for multiple machine vision tasks (e.g., segmentation, and detection) and human vision. 2), a task-specific adapter, that uses the parameter-efficient delta-tuning strategy to stimulate the comprehensive downstream analytical networks for specific machine vision tasks. By using the above two modules, we can optimize the bit-rate costs and improve machine vision performance. In general, our proposed EAC can seamlessly integrate with existing NIC (i.e., Ball'e2018, and Cheng2020) and NVC (i.e., DVC, and FVC) methods. Extensive evaluation on various benchmark datasets (i.e., VOC2007, ILSVRC2012, VOC2012, COCO, UCF101, and DAVIS) shows that our method enhances performance for multiple machine vision tasks while maintaining the quality of human vision.