🤖 AI Summary
While existing EEG-to-image retrieval models demonstrate strong performance on semantically diverse data, their sensitivity to fine-grained visual changes remains unclear. To address this gap, this work introduces EEG-EditBench, the first diagnostic benchmark based on controlled image editing. By systematically manipulating object identity, attributes, background, and object presence, the benchmark comprises 2,137 high-quality edited samples used to evaluate eight representative EEG-based visual decoding models. The results reveal that high retrieval accuracy does not generalize well to scenarios involving subtle visual variations, with model performance weakest at the attribute level. These findings highlight significant limitations in how current models leverage visual information from EEG signals and offer a new perspective for understanding the mechanisms underlying EEG-driven visual representation.
📝 Abstract
Recent EEG-to-image retrieval models have achieved strong performance in identifying viewed images from semantically diverse candidates. Yet such success does not reveal what visual information supports the match. A model may readily identify a cheetah among tools, plants, and vehicles, but can it still distinguish the viewed cheetah from the same scene with the cheetah replaced by a dog? Motivated by this question, we introduce EEG-EditBench, a diagnostic benchmark that examines this question through controlled edits of object identity, attributes, background, and object presence. Built from the 200 THINGS-EEG2 test images, EEG-EditBench contains 2,137 quality-controlled edits and evaluates eight representative EEG visual decoding models. Our results show that strong standard retrieval does not consistently transfer to edit-based evaluation, with fine-grained attribute changes presenting the greatest challenge. EEG-EditBench reveals model behavior hidden by aggregate retrieval accuracy and provides a controlled basis for studying what visual information EEG-image models preserve. The code and complete dataset are publicly available.