Cephalonauts One: A deep fMRI dataset for decoding naturalistic speech in the human brain

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the scarcity of large-scale human fMRI data under natural language stimulation and the absence of standardized decoding benchmarks. To this end, we construct the deepest natural speech fMRI dataset to date, comprising 30 hours of podcast listening data per subject acquired via 3T whole-brain scanning, and achieve multimodal alignment by integrating transcript annotations with feature embeddings. Furthermore, this work pioneers a standardized brain decoding evaluation framework based on audio segment retrieval, providing baseline models and establishing new benchmarks. Experimental results demonstrate that decoding performance consistently improves with increasing training data volume. Collectively, this project provides a high-quality data foundation and a rigorous evaluation paradigm for advancing research in natural speech brain decoding.
📝 Abstract
Cephalonauts One is a whole-brain 3 Tesla (3T) functional magnetic resonance imaging (fMRI) dataset recorded while subjects listened to audio podcasts. Three healthy subjects underwent multiple scanning sessions, each consisting of five 15-minute runs, while listening to podcasts in their native language. With 30 hours of fMRI data per subject, the current release is the deepest available fMRI dataset using naturalistic speech stimuli. The dataset pairs brain activity with the corresponding podcast audio, transcript annotations, and derived stimulus embeddings. Furthermore, we introduce a brain decoding benchmark formulated as audio segment retrieval: given fMRI activity from a held-out session, the decoder must identify the corresponding time-aligned podcast audio segment among candidate segments. We provide standardized splits, evaluation metrics, and baseline decoders for this task. Finally, a scaling analysis shows that decoding performance improves continuously with the amount of training data per subject.
Problem

Research questions and friction points this paper is trying to address.

fMRI dataset
naturalistic speech decoding
brain decoding benchmark
audio segment retrieval
scaling analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

fMRI dataset
brain decoding
naturalistic speech
audio retrieval benchmark
scaling analysis
💼 Related Jobs
No related jobs found.
Antoine Collas
Antoine Collas
INRIA, Université Paris-Saclay
machine learningoptimizationstatistics
L
Louis Jalouzot
UNICOG, CNRS, INSERM, CEA, Paris-Saclay University; LSCP, EHESS, ENS, CNRS, PSL University
G
Géraud Ilinca
Karavela
C
Corentin Caris
Karavela
R
Romain Valabrègue
Centre de NeuroImagerie de Recherche (CENIR), Sorbonne Université, ICM, Paris, France
A
Ahmed Hassayoune
Karavela
D
David Goncalves
Karavela
M
Madeleine Hueber
Karavela
T
Thaddée Delebarre
Karavela
J
Julien Savatovsky
Department of Radiology, Hôpital Fondation Adolphe de Rothschild, Paris, France
C
Clara Fonteneau
Karavela
C
Charles Maussion
Karavela
Bertrand Thirion
Bertrand Thirion
Inria
Machine learningfunctional brain imagingstatistics
Alexis Thual
Alexis Thual
Neurospin, CEA ; Inria Saclay
fMRIoptimal transportbrain decoding