🤖 AI Summary
This study addresses the prohibitive computational costs of pretraining fMRI foundation models by evaluating the necessity of fMRI-specific pretraining. We propose FReD, a framework that leverages a frozen deep compression autoencoder pretrained on natural images to extract fMRI representations, combined with linear probes, late fusion, and shallow Transformers for trait and state prediction. Our findings demonstrate that high performance is achievable without fMRI-specific pretraining, establishing frozen natural image features as an effective baseline for assessing pretraining value. FReD surpasses or matches existing fMRI foundation models across multiple benchmarks while more accurately recovering local signal variations and substantially reducing computational overhead.
📝 Abstract
Foundation models pre-trained on large-scale fMRI datasets have shown strong downstream performance, but at substantial data and computation cost. To investigate how much fMRI-specific pre-training is actually needed for such performance, we introduce FReD, which derives fMRI representations from a frozen Deep Compression AutoEncoder (DCAE) pre-trained exclusively on natural images and pairs them with a task specific readout. For trait prediction, FReD summarizes frame-wise representations by their temporal mean and log-standard deviation and applies linear probing, with late fusion across two normalization schemes. For state prediction, it represents each frame as a single token and models temporal dependencies with a shallow Transformer. Across four resting-state datasets spanning six trait-prediction targets, linear probes on frozen DCAE features generally outperform those on fMRI foundation model representations and remain competitive with fully fine-tuned fMRI foundation models. On three task-fMRI state-prediction tasks, a temporal readout on DCAE features performs comparably to the strongest foundation models evaluated. A Gaussian injection analysis further shows that localized signal changes are recovered more accurately from the frozen DCAE features than from the evaluated foundation-model representations. Together, these results show that strong performance on current fMRI benchmarks is possible without fMRI-specific representation pre-training, making frozen natural-image features as a useful baseline for assessing its added value.