🤖 AI Summary
This study addresses the limitations of existing visual token communication systems, which rely on static strategies incapable of jointly adapting to image content and channel conditions, while token-level optimization struggles to enhance reconstruction quality. To overcome these challenges, this work proposes AdapToC, a framework that, for the first time, co-designs token quantity, unequal error protection, and reliability-aware recovery mechanisms. Specifically, the transmitter adaptively allocates the number of tokens and their corresponding protection levels, while the receiver performs iterative reconstruction by integrating a MaskGIT model with channel state information. Under identical communication overhead, the proposed method achieves a 4.20 dB improvement in peak average PSNR over the strongest static baseline, establishing new state-of-the-art performance in the field.
📝 Abstract
Tokens have become a unified interface for multimodal foundation models, making visual-token communication a natural paradigm for efficient image delivery. However, existing methods typically rely on static policies that cannot jointly adapt to image content and channel conditions. Moreover, their token-level utility objectives do not necessarily translate into improved image reconstruction quality. In this paper, we propose AdapToC, an adaptive, reconstruction-oriented visual-token communication framework. At the transmitter, an adaptive selector jointly models image content, channel state, and communication budget to perform instance-wise resource allocation. Rather than using a fixed token rate and protection policy, it dynamically determines how many tokens should be transmitted and assigns different protection levels according to token importance and current channel conditions. At the receiver, an adaptive MaskGIT receiver incorporates channel reliability into contextual token modeling. It distinguishes tokens with different reliability levels, preserves high-confidence observations, corrects potentially corrupted tokens, and iteratively reconstructs missing content from the received evidence and global visual context. By co-designing token quantity, unequal protection, and reliability-aware recovery for image-level reconstruction quality, AdapToC achieves a peak mean PSNR gain of 4.20 dB over the strongest static baseline under matched communication costs and state-of-the-art performance among the evaluated visual-token communication methods.