Future Video Generation Better Aligns with the Human Visual Cortex than Observed Video

πŸ“… 2026-09-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study investigates the divergent neural response mechanisms in the human visual cortex during the observation versus prediction of future videos, and examines whether generative representations better align with the brain’s predictive properties. By integrating autoregressive and non-autoregressive video diffusion models with fMRI data, this work employs representational similarity analysis and behavioral experiments to systematically compare neural alignment in the visual cortex between future video generation and passive observation. It provides the first evidence that internal representations from future video generation align more closely with human predictive mechanisms than passive observation, an effect specifically localized to higher-order visual cortices. Furthermore, enhancing contributions from higher hierarchical levels significantly increases human preference for generated videos, offering novel insights into the brain’s predictive coding framework.
πŸ“ Abstract
Studying the alignment between the internal representations of vision models and the responses of the visual cortex to the same observed visual stimuli has enabled us to better understand human visual processing. However, studies so far have largely overlooked the fact that the human brain not only processes observed visual stimuli, but also predicts upcoming stimuli based on what has been observed. Accordingly, we hypothesize that internal representations for generating future video frames are better aligned with the predictive nature of human visual processing than representations of the observed video itself. To this end, we compare the alignment between human video-watching fMRI responses in the visual cortex and the internal representations from two types of video diffusion models, an autoregressive (AR) model and its non-AR base model. We first conduct a within-model analysis of the AR video diffusion model and show that the representations for future video generation align better with the visual cortex than the representations of the observed video. We then compare the internal representations of the AR model with those of its non-AR base model and again show that the representations for future video generation align better with the visual cortex than the representations for observed video reconstruction by the base model. Specifically, the alignment of observed video reconstruction is concentrated in lower-order visual cortex, whereas that of future video generation is concentrated in higher-order visual cortex. Finally, we show in a human behavioral experiment that humans prefer videos generated by amplifying the contributions of individual layers that align better with the visual cortex.
Problem

Research questions and friction points this paper is trying to address.

video generation
visual cortex alignment
predictive processing
fMRI
autoregressive model
Innovation

Methods, ideas, or system contributions that make the work stand out.

Video Diffusion Models
Autoregressive Generation
Visual Cortex Alignment
fMRI
Predictive Processing
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
C
Chang-Bae Bang
Department of Psychiatry, Yonsei University College of Medicine; Institute of Behavioral Sciences in Medicine, Yonsei University College of Medicine; Department of Biomedical Systems Informatics, Yonsei University College of Medicine
Hyungjin Chung
Hyungjin Chung
Lead AI Research Scientist, EverEx
Generative ModelsInverse Problems
Byung-Hoon Kim
Byung-Hoon Kim
Yonsei University, College of Medicine
PsychiatryNeuroimagingLarge Multimodal Models