Score
Building and evaluating mappings between computational models and neuroanatomical data by aligning model outputs with brain responses and behavior, combining global and region-specific connectivity, and assessing alignment using neural measurements and human similarity judgments.
This study addresses the limitation of relying solely on prediction accuracy to assess alignment between visual models and human brain responses, as such metrics often obscure which reproducible response dimensions are shared. To overcome this, the authors propose a unified evaluation framework that quantifies the degree to which models or cross-subject brain signals recover reproducible dimensions within a target neural response space, using target-space recovery profiles rather than scalar accuracy measures. Integrating repeated fMRI measurements, cross-run splits, and both brain–brain and model–brain predictive modeling, the approach reveals a low-dimensional, reproducible response structure in early-to-mid-level visual cortex during naturalistic viewing. Notably, despite similar prediction accuracies, pretrained and randomly initialized models exhibit markedly distinct recovery profiles, uncovering alignment differences masked by conventional accuracy metrics.
This study addresses the challenge of comparing high-dimensional neural representations across neuroscience and artificial intelligence: specifically, how to select similarity measures that best reveal functional correspondences and divergences. We systematically evaluate eight mainstream representational similarity metrics—including linear CKA, Procrustes distance, CCA, inner-product kernel, and nearest-neighbor alignment—against behavioral functional alignment (e.g., recognition accuracy, generalization, robustness) as a ground-truth benchmark. Our evaluation spans both biological neural data and artificial neural network models. Results show that geometry-sensitive metrics—particularly linear CKA and Procrustes distance—consistently outperform predictive metrics, achieving superior alignment with human behavioral performance and effectively distinguishing trained versus untrained models. In contrast, linear predictivity exhibits only moderate behavioral alignment. This work establishes the first behavior-driven representational similarity benchmark, providing a principled, cross-domain methodology for mechanistic interpretation and comparative analysis of neural computation.
This work investigates the alignment between neural network representations and human psychological representations—operationalized via behavioral similarity judgments. Using representational similarity analysis (RSA), linear alignment, and multi-task cognitive evaluation, we systematically assess the influence of model scale, architecture, training data, and objective functions across three behavioral datasets. We find that dataset selection and loss function design are primary determinants of alignment quality, whereas model scaling yields negligible improvement. We propose a cross-dataset linear transformation method that significantly enhances generalization of alignment across domains. Furthermore, we observe that networks robustly encode natural categories (e.g., food, animals) but exhibit weak representation of culturally or abstractly grounded concepts (e.g., “royalty,” “sports”). Our core contributions are: (1) identifying key drivers of representational alignment—particularly data and optimization choices over architectural or scale factors; and (2) introducing a transferable, linear calibration strategy for improving cross-dataset alignment fidelity.
Systematic comparison between artificial vision models and human brain representational spaces remains methodologically fragmented and lacks standardized, reproducible frameworks. Method: The authors developed an open-source Python toolbox that unifies over 600 cross-modal pretrained models with major neuroimaging datasets (e.g., NSD, THINGS), enabling end-to-end model–brain comparison—from model invocation and neural data acquisition to feature extraction, representational similarity analysis (RSA), and visualization. The toolbox implements a standardized neural representational alignment pipeline, supporting advanced neuroencoding techniques including searchlight analysis, linear encoding modeling, and variance decomposition. Contribution/Results: This framework substantially improves reproducibility, flexibility, and scalability in model–brain comparative research. It has been adopted by multiple cognitive neuroscience laboratories and is advancing standardization at the intersection of computational neuroscience and artificial intelligence.
This study investigates the alignment mechanisms between neural network representations and human visual learning in few-shot image understanding. Method: We systematically evaluate generalization behaviors of 86 pretrained models on continuous relational reasoning and natural image classification tasks, introducing the first quantitative measure of cross-task consistency between model representations and human cognitive trajectories. Our approach integrates representational similarity analysis, intrinsic dimension estimation, cognitive-modeling–driven evaluation, and multimodal contrastive learning assessment. Results: Multimodal contrastive learning emerges as the strongest predictor of human few-shot generalization—significantly outperforming conventional metrics such as parameter count or training data scale. Pretrained models serve as effective sources of cognitive representations. The proposed evaluation paradigm establishes a generalizable, ecologically valid framework for cross-species intelligence modeling, advancing the study of human-aligned artificial perception.
This work addresses the limited generalization performance in cross-subject brain functional decoding caused by inter-individual variability in neural responses. To overcome this challenge, the authors propose SpectralOT, a novel method that, for the first time, integrates spectral features of the Laplace–Beltrami operator into functional data and leverages optimal transport theory to construct a geometry-aware whole-brain alignment framework. By explicitly incorporating cortical geometric structure during functional alignment, the approach enhances computational efficiency while preserving anatomical consistency. Experimental results demonstrate that SpectralOT significantly improves the generalization capability of cross-subject decoding models, offering a new paradigm for high-precision brain functional analysis.
Existing activation alignment methods struggle to capture differences in the sensitivity of neural representations to local stimulus perturbations and thus fail to reflect how systems leverage local evidence for discrimination. This work proposes a novel analytical framework based on locally decodable information, integrating Fisher information, pullback metrics, and log-spectral distances on the SPD manifold to construct the Spectral Riemannian Alignment Score (S-RAS). S-RAS provides, for the first time, a minimal, dataset-level summary of neural representational sensitivity from the perspective of local discriminative tasks, with guaranteed multiplicative consistency. The method successfully aligns corresponding layers across independently trained networks, enables transferable class-conditional probing, reveals representational differences between standard and robustly trained models, and uncovers stimulus coordinate family effects in mouse visual cortex.
This work proposes the Neural Functional Alignment Space (NFAS) to address the challenge of uniformly evaluating representations across diverse neural architectures, which existing brain-referenced approaches struggle with due to their reliance on static layers or task-specific alignment. NFAS leverages Dynamic Mode Decomposition (DMD) to extract the evolutionary trajectory of stimulus representations along network depth and projects these trajectories into a biologically anchored coordinate system constructed from multimodal neural responses. This enables a unified, cross-modal, and cross-architectural characterization of model functionality. The authors introduce the Signal-Noise Consistency Index (SNCI) to quantify alignment quality. Experiments across 45 pretrained models reveal that representations in NFAS cluster by modality and exhibit cross-modal convergence within integrative cortical systems, demonstrating that representational dynamics provide a general foundation for functional model evaluation.
This study addresses a critical gap in brain alignment research, which has largely focused on language comprehension or passive visual tasks, by systematically investigating the alignment between foundational models and human brain activity during naturalistic interaction. For the first time, the authors apply a vision-language model (VLM) and a large action model (LAM) to analyze fMRI data recorded while participants played Atari-style games. Using voxel-wise encoding and variance partitioning, they assess how internal model representations align with neural responses under varying prompting conditions. Results demonstrate that both VLM and LAM significantly outperform reinforcement learning baselines. Prompt effects intensify along the cortical hierarchy, peaking in frontoparietal and motor planning regions. Moreover, LAM representations preferentially align with action-related areas, whereas VLM exhibits more symmetric cross-modal alignment across sensory and cognitive systems.