🤖 AI Summary
Existing brain signal decoding methods typically align only the global embeddings of Vision Transformers (ViTs), neglecting the rich visual information encoded in intermediate patch representations. This work proposes a Residual Patch Adapter, a lightweight and modular design that leverages the full set of patch tokens across all intermediate ViT layers to achieve multi-granularity brain-image alignment. Our analysis demonstrates that retaining all patches significantly outperforms pooling or masking strategies, revealing the critical role of synergistic low- and high-level features in cross-modal alignment. Evaluated on the THINGS dataset, the proposed method achieves state-of-the-art performance, yielding substantial improvements in cross-subject Top-1 accuracy. These results validate both the effectiveness and generalization capability of our approach for fine-grained neural decoding.
📝 Abstract
Most existing MEG and EEG (M/EEG) visual decoding methods align brain signals with a single global embedding extracted from a pretrained visual encoder, leaving open whether intermediate patch representations, which preserve richer and more granular rich visual information, can improve representation learning. To address this question, we introduce the Residual Patch Adapter (RPA), a lightweight, modular adapter that leverages all patch tokens from an intermediate layer of a ViT visual encoder for alignment. Through extensive ablation analyses, we first show that pooling or masking patch tokens degrades the learned representation, demonstrating that retaining the full set of patch tokens is important for EEG alignment, while the CLS token provides little unique information. We then use a series of six quantitative feature analyses to show that both higher-level semantics and lower-level visual features, including color and texture, are essential for this EEG-to-image alignment. Under current protocols, our system achieves Top-1 accuracies of 95.4\% within-subject and 35.5\% cross-subject on THINGS-EEG2, and 65.2\% and 6.7\%, respectively, on THINGS-MEG, achieving state-of-the-art (SOTA) performance across both datasets. Evaluations with alternative brain encoders, including pretrained EEG foundation models, demonstrate that the approach extends beyond the projection-based EEG encoder. Furthermore, we provide a plug-and-play interface that allows RPA to be replaced by convolution, attention, or ConvNeXt alternatives. Together, these findings provide significant insight into M/EEG-to-image representation learning by establishing design principles for leveraging the latent space of visual encoders, and open new directions for brain--image alignment and non-invasive brain--computer interface (BCI).