🤖 AI Summary
This study addresses the challenges of high visual data transmission latency and substantial continuous feature encoding overhead in on-device multimodal reasoning by proposing a query-guided, task-oriented feature compression method. Specifically, the approach employs residual vector quantization (RVQ) to discretize continuous visual features into compact index sequences, thereby reducing representation costs. Furthermore, it designs a query-relevance aggregation mechanism to precisely extract task-critical information and introduces an error-compensation adapter to mitigate quantization distortion. Experimental results demonstrate that the proposed method maintains comparable task performance while reducing the visual transmission payload by 53.6% relative to baselines. Consequently, this work significantly optimizes end-to-end inference latency in bandwidth-constrained scenarios.
📝 Abstract
Large multimodal models (LMMs) support diverse visual understanding and reasoning tasks but are often impractical to run entirely on resource-constrained devices. Device-edge co-inference reduces device computation, yet transmitting visual data over bandwidth-limited uplinks can introduce substantial delay. Task-oriented feature compression (TOFC) reduces the payload through feature aggregation and entropy coding. However, continuous-feature coding remains costly, and query-agnostic aggregation may discard task-relevant local evidence. We propose query-guided task-oriented feature compression (Q-TOFC) for device-edge multimodal inference. Q-TOFC employs residual vector quantization (RVQ) to encode each merged feature as a compact sequence of codebook indices, reducing its representation cost and allowing more features to be transmitted. It further incorporates query relevance into feature aggregation and uses a quantization error compensation adapter to mitigate the distortion introduced by discrete quantization. Experiments on seven multimodal benchmarks show that Q-TOFC reduces the visual payload by 53.6% relative to TOFC while maintaining comparable average normalized task performance. End-to-end latency evaluations further demonstrate lower latency under bandwidth-constrained uplinks.