VQ-LIC: Shared Vector-Quantized Learned Image Compression on a Resource-Constrained FPGA

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the memory-compute imbalance in deploying learned image compression on resource-constrained FPGAs by proposing an asymmetric edge-cloud codec architecture. The edge device reuses depthwise/pointwise (DW/PW) engines to execute INT8 transforms and vector quantization (VQ), while the cloud handles reconstruction. A core innovation is the first mapping of VQ onto PW engines, eliminating dedicated array overhead, complemented by an RTL-based deterministic latency model that guides architecture search. Combined with hardware-aware optimizations such as multi-codebook VQ and post-training pruning, the proposed method achieves 47.98 FPS and 42.84 mJ/frame energy efficiency on a Zynq-7020 FPGA using minimal DSP resources, yielding rate-distortion performance superior to larger-scale models.
📝 Abstract
Learned image compression (LIC) is hard to deploy on severely resource-constrained FPGAs, since how fast it actually runs depends not just on arithmetic count, but also on memory traffic, imbalance between different operations, and how the hardware batches its work. We present VQ-LIC, an asymmetric edge-cloud codec in which a compact INT8 depthwise (DW)-pointwise (PW) analysis transform and multi-codebook vector quantization (VQ) run at the edge on a reusable DW/PW engine pair, while reconstruction is handled by a larger cloud decoder. Since VQ codeword matching is expressible as a dot product, it is mapped directly onto the same PW engine, removing the need for a separate VQ compute array, to our knowledge a first for FPGA LIC. A novel latency model, derived from deterministic RTL cycle counts of an FPGA's read, DW, PW, and write costs, predicts an analysis transform's per-block latency; since VQ shares the same PW datapath, the model applies to VQ as well. Validated directly against silicon, the model predicts deployed analysis and VQ latency within 0.26\% and 0.05\%, and guides the selection of a three-block $16$-$48$-$64$ transform. Post-training codebook reduction then cuts VQ arithmetic and codebook storage by $4\times$ and shrinks the fixed-width latent representation. On a 220-DSP Zynq-7020, VQ-LIC's mid-rate preset reaches 0.1398 bits per pixel at 28.69 dB PSNR and 13.06 dB MS-SSIM on CLIC~2017, outperforming a similarly sized neural encoder and reaching a rate-distortion range comparable to a codec three orders of magnitude larger. The complete 0.1945-kMAC/pixel analysis-VQ pipeline runs at 47.98 frames per second and 42.84 mJ per frame on silicon, using an order of magnitude fewer DSPs than comparable FPGA LIC accelerators while achieving lower bitrate, higher throughput, and lower energy per frame at a modest PSNR tradeoff.
Problem

Research questions and friction points this paper is trying to address.

Learned Image Compression
Resource-Constrained FPGA
Memory Traffic
Hardware Deployment
Edge Computing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Learned Image Compression
Vector Quantization
FPGA Acceleration
Edge-Cloud Codec
Latency Modeling
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Muhammad Fahd Ibrahim Bhatti
School of Science and Engineering (SBASSE), Lahore University of Management Sciences (LUMS), Lahore, Pakistan
A
Abdullah Bin Faisal
School of Science and Engineering (SBASSE), Lahore University of Management Sciences (LUMS), Lahore, Pakistan
A
Ahsan Usman
School of Science and Engineering (SBASSE), Lahore University of Management Sciences (LUMS), Lahore, Pakistan
Naveed Anwar Bhatti
Naveed Anwar Bhatti
Assistant Professor, LUMS, Lahore
Cyber Physical SystemsInternet of ThingsEmbedded SystemsWireless Sensor NetworksSecurity
M
Muhammad Ali Siddiqi
School of Science and Engineering (SBASSE), Lahore University of Management Sciences (LUMS), Lahore, Pakistan