🤖 AI Summary
This study achieves end-to-end semantic decoding of electrocorticographic (ECoG) signals under few-shot conditions (fewer than 50 samples per class), directly predicting the category of dynamic video stimuli from high-gamma band (80–150 Hz) neural activity. The proposed approach employs a Transformer encoder to process ECoG signals within a 900 ms post-stimulus time window and incorporates mixup data augmentation to enhance generalization. Notably, the model operates without handcrafted features, maintaining high performance while offering strong interpretability. It identifies key contributing regions—including visual areas V2–V4, the ventral visual stream, the MT+ complex, and lateral temporal cortex—whose involvement aligns closely with established neuroscientific understanding of visual processing.
📝 Abstract
ECoG-based visual semantic decoding enables inference of semantic interpretation of visual perception from complex, noisy brain activity. This study examines the feasibility of visual semantic decoding using an end-to-end deep learning framework using electrocorticography (ECoG). Specifically, the decoding task is to predict visual categories from video stimuli using time-series neural inputs. A previously collected ECoG dataset from participants ($n=17$) with drug-resistant epilepsy is used for analysis. With fewer than 50 training samples per visual category, this study evaluates multiple deep learning approaches, artificial neural network architectures, and frequency-band filtered inputs. The best-performing approach is analyzed to shed light on the discriminative information it relies on across spectral, temporal, and cortical dimensions. The selected decoding system uses mixup augmentation, a Transformer-based encoder, and high-gamma (80-150 Hz) inputs with a 900 ms post-stimulus window. Further analysis shows that early visual cortex (V2-V4), ventral stream visual cortex, MT+ complex with neighbouring visual areas, and lateral temporal cortex contributed substantially to decoding performance. This study demonstrates that an end-to-end deep learning framework can yield promising decoding performance from dynamic visual stimuli without handcrafted features, while the model behavior remains interpretable through spectral, temporal, and cortical dimensions, which are broadly consistent with established neuroscience knowledge.