π€ AI Summary
This study addresses the challenge in cross-subject EEG-based image retrieval where source-domain neural representations struggle to align with the visual space of unseen subjects, proposing a novel framework optimized from the visual target side. The method transforms perceptual encoder patch grids into compact visual views, which are aggregated via a block-structured content-dependent router and jointly trained with an EEG encoder using contrastive learning and MMD regularization. Additionally, a training-free representation refinement strategy is designed to achieve cross-domain alignment of frozen embeddings without updating the encoders. Evaluated on the THINGS-EEG2 dataset, the proposed approach attains a Top-1 accuracy of 48.1%, outperforming the strongest baseline by 18.5% and significantly improving retrieval performance across all held-out subjects.
π Abstract
Cross-subject EEG-to-image retrieval requires a neural represen- tation trained on source subjects to remain aligned with a visual embedding space for an unseen subject. Whereas existing methods primarily focus on the EEG side, we address this problem from the perspective of the visual target. Our approach preserves the spatial information of the Perception Encoder, converts its patch grid into a compact set of learned visual views, and aggregates them for each image with a block-structured, content-dependent router. The target is learned jointly with the EEG encoder through contrastive learning with MMD regularization across source subjects. For deployment, we propose a training-free representation refinement that aligns frozen embeddings without updating either encoder. Under leave- one-subject-out evaluation on THINGS-EEG2, the structured target achieves 35.3%/65.6% Top-1/Top-5 accuracy, the best among com- pared methods. Refinement raises this to 48.1%/77.1%, an 18.5% Top-1 gain over the strongest compared method, improving all ten held-out subjects.