Samples, Sources, Space: Decomposing Data Scale in Spatially Structured Representation Learning of Human Brain Microarchitecture

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of conventional data scaling research that overlooks disparities in sample size, source diversity, and spatial coverage within spatially structured data by reformulating data scaling as a resource allocation problem. Methodologically, it establishes these three dimensions as independent scaling axes for the first time, employing a spatial proximity-supervised contrastive learning model to process 11.6 million whole-brain tissue sections. Results demonstrate that representation performance improves with increased sample size, spatial coverage, computational resources, and model capacity. Furthermore, the findings reveal that under fixed budgets, merely expanding data sources yields no benefit, and cross-subject generalization is constrained by biological variation rather than driven by the number of sources.
📝 Abstract
Scaling studies typically represent training data by a single count of samples. For hierarchically and spatially structured data, however, the same number of samples can be drawn from few or many sources and distributed differently across the underlying domain. We therefore study data scaling as an allocation problem, separating unique sample count, source diversity, and spatial coverage. We study this decomposition in microscopic whole-brain histology, where a source is an individual brain, and a sample is an image patch at a specific spatial location. Across 93 controlled pretraining runs of a contrastive model that uses spatial proximity for supervision, we vary data allocation, compute, and model capacity over 11.6 million spatially anchored image patches from 21 human brains. Performance improves with more unique samples, broader spatial coverage, additional compute, and larger model capacity. At fixed sample count, distributing samples across one to 18 subjects produces no detectable improvement, even though representations generalize substantially better to subjects encountered during pretraining. Inter-subject variation therefore strongly affects generalization, but additional subjects provide no benefit when a fixed sample budget is distributed across more sources. These results establish sample count, source diversity, and spatial coverage as distinct axes of data scaling in spatially structured representation learning.
Problem

Research questions and friction points this paper is trying to address.

representation learning
data scaling
spatial structure
source diversity
brain microarchitecture
Innovation

Methods, ideas, or system contributions that make the work stand out.

data scaling decomposition
spatially structured representation learning
contrastive learning
brain microarchitecture
source diversity
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Christian Schiffer
Institute of Neuroscience and Medicine (INM-1), Research Centre Jülich, Jülich, Germany; Helmholtz AI, Research Centre Jülich, Jülich, Germany
M
Mathis Bode
Jülich Supercomputing Centre, Research Centre Jülich, Jülich, Germany
Thomas Lippert
Thomas Lippert
Professor for Computer Science, University Frankfurt and FIAS, Director Jülich Supercomputing Centre
Lattice quantum chromodynamicsnumerical algorithmsneurosciencecomputer architecturesquantum computing
K
Katrin Amunts
Institute of Neuroscience and Medicine (INM-1), Research Centre Jülich, Jülich, Germany; Cécile & Oscar Vogt Institute for Brain Research, University Hospital Düsseldorf, Düsseldorf, Germany
T
Timo Dickscheid
Institute of Neuroscience and Medicine (INM-1), Research Centre Jülich, Jülich, Germany; Helmholtz AI, Research Centre Jülich, Jülich, Germany; Computer Vision, Institute for Computational Visualistics, University of Koblenz, Koblenz, Germany