Efficient Masked Image Compression with Position-Indexed Self-Attention

📅 2025-04-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Current semantic mask-based image compression methods still encode and process the entire image, leading to redundant computation on background regions—even when their pixels are zeroed out or omitted from transmission. To address this, we propose the first mask-driven sparse coding-decoding framework explicitly designed for high-level vision tasks. Our approach leverages positional-indexed self-attention to extract and reconstruct features exclusively within semantic mask-visible regions, eliminating dependence on full-image inputs. Coupled with sparse feature modeling and a mask-guided lightweight coding-decoding architecture, it reduces computational overhead at the source. Experiments demonstrate that our method significantly lowers encoding complexity and bit-rate while preserving downstream task performance (e.g., object detection and semantic segmentation), outperforming existing semantic-structured compression approaches.

Technology Category

Computer Vision: SegmentationMachine Learning: Learning on the Edge & Model CompressionCognitive Modeling & Cognitive Systems: Neural Spike Coding

Application Category

Search and Retrieval-Augmented AI: Efficiency and scalability of Web search enginesEconomics, Online Markets and Human Computation: Architectures and workflows that use LLMs for crowd workSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
In recent years, image compression for high-level vision tasks has attracted considerable attention from researchers. Given that object information in images plays a far more crucial role in downstream tasks than background information, some studies have proposed semantically structuring the bitstream to selectively transmit and reconstruct only the information required by these tasks. However, such methods structure the bitstream after encoding, meaning that the coding process still relies on the entire image, even though much of the encoded information will not be transmitted. This leads to redundant computations. Traditional image compression methods require a two-dimensional image as input, and even if the unimportant regions of the image are set to zero by applying a semantic mask, these regions still participate in subsequent computations as part of the image. To address such limitations, we propose an image compression method based on a position-indexed self-attention mechanism that encodes and decodes only the visible parts of the masked image. Compared to existing semantic-structured compression methods, our approach can significantly reduce computational costs.
Problem

Research questions and friction points this paper is trying to address.

Selective image compression for high-level vision tasks
Reducing redundant computations in masked image encoding
Position-indexed self-attention for efficient partial image processing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Position-indexed self-attention for masked image compression
Encodes and decodes only visible parts of masked image
Significantly reduces computational costs in compression
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Chengjie Dai
Zhejiang University
T
Tiantian Song
University College London
H
Hui Tang
Zhejiang University
F
Fangdong Chen
Hikvision Research Institute
B
Bowei Yang
Zhejiang University
Guanghua Song
Guanghua Song
浙江大学
machine learning