π€ AI Summary
This study addresses the limitation of existing video encoding models that overlook the spatiotemporal structure of neural responses, hindering the prediction of dynamic visual processing mechanisms in the brain under naturalistic video conditions. We propose the first video-oriented joint spatiotemporal cross-attention framework, which leverages features from the self-supervised model V-JEPA-2 for spatiotemporally joint routing to construct an interpretable fMRI encoding model simulating the cortexβs adaptive weighting of dynamic information. This model significantly improves brain response prediction accuracy in higher-order visual areas and generates interpretable attention maps capable of tracking moving objects. Furthermore, it reveals dynamically shifting attention allocation patterns within motion perception and semantic selection networks as objects move, thereby elucidating the computational mechanisms underlying dynamic visual perception.
π Abstract
Understanding how the brain parses actions and events from time-varying natural inputs is a central challenge in neuroscience. Recent work has used deep neural network (DNN) models to build stimulus-computable fMRI encoding models that predict single-voxel responses to complex natural videos. However, the majority of video-computable encoding models predict responses using simple linear mappings from model tokens, overlooking the spatiotemporal structure shared by video representations and neural responses. Recent cross-attention encoding models address this limitation for static images, enabling flexible stimulus-dependent weighting of image content across space. Here, we extend this framework to naturalistic video, using per-parcel cross-attention to dynamically route features from a self-supervised video model (V-JEPA-2) across both space and time, fitting this model to fMRI responses to short video clips. We compare joint spatiotemporal attention with factorized and selectively constrained alternatives, and find that joint routing improves predictions of brain responses to held-out videos across higher visual regions, most consistently in lateral and dorsal visual areas associated with dynamic motion perception. Moreover, our method provides interpretable, stimulus-specific attention maps that dynamically follow moving objects, revealing which locations and temporal moments contribute to each neural response. We further show that attention maps from parcels in different category-selective networks (face-, body-, scene-selective) differentially weight content in accordance with expected semantic selectivity. Together, this work provides a new computational framework for understanding how visual information is adaptively weighted by cortical populations during dynamic visual perception.