Embedding Compression Distortion in Video Coding for Machines

📅 2025-03-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing video codecs are optimized for human visual perception, neglecting the impact of compression artifacts on machine vision tasks—leading to degraded downstream AI performance. This paper proposes CDRE, a framework that models and compresses “compression-sensitive distortion” in the feature domain as a learnable, transmittable machine-perception embedding. Methodologically, CDRE introduces: (1) the first distortion-driven differentiable embedding paradigm; (2) a co-designed architecture integrating a compression-sensitive feature extractor with a lightweight distortion codec (comprising quantization and entropy modeling); and (3) progressive embedding adaptation and rate–task joint optimization during training. Evaluated on H.264, H.265, and AV1 encoders across object detection and action recognition tasks, CDRE achieves an average mAP gain of 3.2%, with marginal overhead—less than 0.5% bitrate increase and under 1% additional model parameters.

Technology Category

Machine Learning: Learning on the Edge & Model CompressionComputer Vision: Diffusion Models for VisionData Mining & Knowledge Management: Data Compression

Application Category

Graph Algorithms and Modeling for the Web: Graph embeddings and representation learning for Web-related graphsSearch and Retrieval-Augmented AI: Vertical and domain-specific searchUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and ranking
📝 Abstract
Currently, video transmission serves not only the Human Visual System (HVS) for viewing but also machine perception for analysis. However, existing codecs are primarily optimized for pixel-domain and HVS-perception metrics rather than the needs of machine vision tasks. To address this issue, we propose a Compression Distortion Representation Embedding (CDRE) framework, which extracts machine-perception-related distortion representation and embeds it into downstream models, addressing the information lost during compression and improving task performance. Specifically, to better analyze the machine-perception-related distortion, we design a compression-sensitive extractor that identifies compression degradation in the feature domain. For efficient transmission, a lightweight distortion codec is introduced to compress the distortion information into a compact representation. Subsequently, the representation is progressively embedded into the downstream model, enabling it to be better informed about compression degradation and enhancing performance. Experiments across various codecs and downstream tasks demonstrate that our framework can effectively boost the rate-task performance of existing codecs with minimal overhead in terms of bitrate, execution time, and number of parameters. Our codes and supplementary materials are released in https://github.com/Ws-Syx/CDRE/.
Problem

Research questions and friction points this paper is trying to address.

Optimize video coding for machine vision tasks
Address compression distortion in machine perception
Embed distortion representation to improve task performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Embed machine-perception distortion into models
Design compression-sensitive feature extractor
Lightweight codec for efficient distortion transmission
🔎 Similar Papers
No similar papers found.