🤖 AI Summary
This study addresses the challenge of establishing cross-subject correspondence for superficial white matter short fibers due to individual anatomical variability. We propose a multimodal clustering framework based on vision-language models, which innovatively converts multi-parcel cortical endpoints into textual encodings and integrates trajectory geometry, cortical context, and shape information into a unified embedding representation. Leveraging pretrained large language models and ultra-high-resolution diffusion MRI data, this approach achieves highly accurate short fiber organization at the population level. Experimental results demonstrate that the model significantly enhances cluster consistency, successfully recovering 96.7% of learned clusters (5,000 in total) in unseen subjects, thereby validating its exceptional generalization capability.
📝 Abstract
The superficial white matter (SWM), a critical brain region for cognition across the lifespan and brain disease, contains abundant short-range association fibers whose organization remains incompletely characterized, in part because the short trajectories and highly variable cortical folding make correspondence across individuals challenging. Anatomically corresponding connections may vary in spatial location across individuals and therefore may not be adequately defined by geometric proximity alone. We introduce FiberGeoText (FGT), a vision-language model (VLM) for organizing short-range superficial white matter (SWM) streamlines reconstructed from ultra-high-resolution diffusion MRI into population-level clusters. FGT jointly represents three complementary properties of each streamline: its three-dimensional trajectory, its cortical anatomical context, and its shape. Cortical endpoint information from multiple parcellation schemes is expressed as text and encoded using a pretrained large language model (LLM), enabling heterogeneous anatomical descriptions to contribute to a common continuous representation. We evaluated FGT on acquired submillimeter 0.76 mm diffusion MRI data. Compared with state-of-the-art (SOTA) methods, FGT produced substantially greater cortical parcel coherence, within-cluster shape consistency, cluster-size consistency, and cross-subject correspondence. The trained model also generalizes well to unseen subjects with an average of 96.7% of the 5,000 learned clusters recovered, and high consistency of cluster structure between training and testing data. Together, these findings demonstrate that integrating geometric, anatomical, and shape information by learning multimodal deep embeddings with a VLM model enables robust learning of population-consistent SWM organization despite interindividual anatomical variability.