Task-Oriented Visual Feature Compression via Residual Vector Quantization for Device-Edge Multimodal Inference

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of high visual data transmission latency and substantial continuous feature encoding overhead in on-device multimodal reasoning by proposing a query-guided, task-oriented feature compression method. Specifically, the approach employs residual vector quantization (RVQ) to discretize continuous visual features into compact index sequences, thereby reducing representation costs. Furthermore, it designs a query-relevance aggregation mechanism to precisely extract task-critical information and introduces an error-compensation adapter to mitigate quantization distortion. Experimental results demonstrate that the proposed method maintains comparable task performance while reducing the visual transmission payload by 53.6% relative to baselines. Consequently, this work significantly optimizes end-to-end inference latency in bandwidth-constrained scenarios.
📝 Abstract
Large multimodal models (LMMs) support diverse visual understanding and reasoning tasks but are often impractical to run entirely on resource-constrained devices. Device-edge co-inference reduces device computation, yet transmitting visual data over bandwidth-limited uplinks can introduce substantial delay. Task-oriented feature compression (TOFC) reduces the payload through feature aggregation and entropy coding. However, continuous-feature coding remains costly, and query-agnostic aggregation may discard task-relevant local evidence. We propose query-guided task-oriented feature compression (Q-TOFC) for device-edge multimodal inference. Q-TOFC employs residual vector quantization (RVQ) to encode each merged feature as a compact sequence of codebook indices, reducing its representation cost and allowing more features to be transmitted. It further incorporates query relevance into feature aggregation and uses a quantization error compensation adapter to mitigate the distortion introduced by discrete quantization. Experiments on seven multimodal benchmarks show that Q-TOFC reduces the visual payload by 53.6% relative to TOFC while maintaining comparable average normalized task performance. End-to-end latency evaluations further demonstrate lower latency under bandwidth-constrained uplinks.
Problem

Research questions and friction points this paper is trying to address.

Task-oriented feature compression
Device-edge co-inference
Multimodal inference
Visual payload reduction
Feature aggregation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Task-Oriented Feature Compression
Residual Vector Quantization
Device-Edge Co-Inference
Query-Guided Aggregation
Quantization Error Compensation
🔎 Similar Papers
L
Luning Pang
Department of Electronic Engineering, and Beijing National Research Center for Information Science and Technology, Tsinghua University, Beijing 100084, China; and Institute of Artificial Intelligence (TeleAI), China Telecom, Beijing 100033, China
Cheng Yuan
Cheng Yuan
Associate Professor, School of Mathematics and Statistics, Central China Normal University
Computational PhysicsDeep Learning
J
Jiawei Shao
Institute of Artificial Intelligence (TeleAI), China Telecom, Beijing 100033, China
M
Mingtao Huang
Department of Electronic Engineering, and Beijing National Research Center for Information Science and Technology, Tsinghua University, Beijing 100084, China
Yuan Shen
Yuan Shen
Professor, EE, Tsinghua University
LocalizationCommunication and SensingMulti-agent Systems