Score
Design and train models that encode anatomical structure into compact, interpretable latent representations from imaging or volumetric anatomical data, typically using variational autoencoders and their semi‑supervised variants. These representations are constructed to reconstruct or generate aligned segmentation masks, be conditioned on clinical priors, and to disentangle static anatomy from temporal or dynamic factors for downstream analysis and tasks.
This work addresses key clinical challenges in medical point cloud analysis—including data scarcity, substantial inter-patient anatomical variability, and insufficient model interpretability and robustness. We systematically review deep learning–based 3D shape analysis advances from 2021 to 2025, focusing on three core tasks: registration, reconstruction, and anatomical variation modeling. To overcome these challenges, we propose a novel paradigm integrating hybrid geometric representations, large-scale self-supervised pretraining, and generative modeling—balancing structural fidelity with semantic interpretability. Our approach is rigorously evaluated on benchmark medical datasets (e.g., FAUST, OAI, ShapeNet-Med) under a unified assessment framework. The study establishes the first comprehensive, clinically oriented reference framework for medical point cloud shape learning, explicitly identifying technical bottlenecks and translational pathways. It provides both theoretical foundations and practical guidelines for developing trustworthy, generalizable anatomical models suitable for real-world clinical deployment.
Early detection of Alzheimer’s disease (AD) remains challenging due to the difficulty in identifying interpretable, imaging-based biomarkers from structural MRI. Method: We propose an interpretable unsupervised deep learning framework centered on a lightweight 3D convolutional autoencoder that learns compact latent representations of MRI data. Multi-stage dimensionality reduction—integrating PCA and UMAP—is combined with the AAL brain atlas for neuroanatomically grounded visualization. Critically, we introduce Latent Region Correlation Profiling (LRCP), a novel analytical framework that jointly applies SHAP-based regression and assumption-free statistical testing to quantify region-specific contributions to cognitive status variability. Results: Our lightweight model robustly captures AD-progressive anatomical patterns without supervision. LRCP substantially enhances both clinical interpretability and neuroanatomical fidelity of latent features, enabling precise identification of AD-relevant brain regions. This framework establishes a new paradigm for unsupervised AD biomarker discovery—balancing discriminative power with biological interpretability.
This work addresses the scarcity of high-quality brain segmentation masks in non-contrast CT neuroimaging, particularly due to the high annotation cost and substantial variability of ischemic infarcts. To tackle this challenge, the authors propose an anatomy-preserving generative framework that, for the first time, integrates a diffusion model into the latent space of a variational autoencoder (VAE) trained on segmentation masks. The method enables unconditional generation of multi-class brain tissue masks containing ischemic infarcts and supports coarse control over lesion presence via binary prompts. By leveraging a frozen VAE decoder to reconstruct masks, the approach effectively preserves global anatomical structure, discrete semantic labels, and realistic pathological variations while avoiding common structural artifacts associated with pixel-level generative models.
Extracting biologically meaningful and clinically interpretable representations from high-dimensional neuroimaging data (e.g., MRI/PET) remains challenging due to inherent complexity and limited interpretability of latent features. Method: This study systematically reviews and empirically evaluates generative latent-variable models—including autoencoders, GANs, and latent diffusion models (LDMs)—across two complementary pathways: clinical neuroimaging and computational neuroscience. It pioneers the integration of predictive coding theory with deep generative modeling to establish a multimodal alignment and interpretable latent-space analysis framework, accompanied by a cross-model performance evaluation protocol. Contribution/Results: The work delineates the applicability boundaries of each model class for Alzheimer’s disease and Parkinson’s disease subtyping, longitudinal tracking, and brain-age estimation. It significantly enhances the biological interpretability and clinical transferability of learned latent representations, providing a methodological foundation for interpretable brain-computational modeling.
In virtual imaging trials, generating anatomically accurate, clinically relevant patient-specific phantoms with controllable population-level anatomical variations remains challenging. Method: We propose the first implicit neural representation framework for editable anatomical modeling, integrating geometry-prior-guided implicit surface reconstruction, disentangled latent space learning, and topology-adaptive deformation—enabling fine-grained, target-specific morphological editing of topologically variable organs (e.g., thyroid). Contribution/Results: Our approach is the first to achieve explicit shape–topology disentanglement in anatomical implicit neural representations, supporting clinically interpretable, parameterized editing. Quantitative and qualitative evaluations demonstrate state-of-the-art performance in reconstruction accuracy and anatomical plausibility. Generated phantoms exhibit high fidelity, clinical interpretability, and strong controllability—facilitating reproducible, patient-population-aware virtual imaging studies.
This study investigates the structure and information content of latent representations in 3D brain MRI generative models, with a focus on their efficacy for clinical discrimination of Down syndrome. Employing various variational autoencoder (VAE) architectures, we compress 3D brain MRIs into compact latent codes that enable high-fidelity reconstruction while supporting downstream multitask analysis. Through principal component analysis visualization and systematic evaluation, we demonstrate that the learned latent space clearly clusters individuals with Down syndrome and neurotypical controls, exhibiting strong discriminative power and interpretability. These findings validate the potential of such latent representations for clinical neuroimaging applications and offer a novel approach to disease representation learning based on generative models.
研究对比了解剖特征与学习特征在脑MRI分析中的效果,提出结合解剖信息预训练的新方法,提升了生物年龄估计的准确性。
This work addresses the limitations of traditional statistical shape modeling, which relies on dense annotations and fixed latent representations, thereby struggling to flexibly capture complex anatomical variations. The authors propose MorphoFlow, a framework that learns compact probabilistic shape representations from only sparse surface annotations. MorphoFlow integrates neural implicit representations, a self-decoder architecture, and autoregressive normalizing flows, augmented with an adaptive latent correlation weighting mechanism. This mechanism leverages a sparsity-inducing prior to automatically modulate the contribution of each latent dimension to anatomical variability, eliminating the need for manual hyperparameter tuning. The method enables high-resolution 3D shape generation and uncertainty quantification. Evaluated on lumbar spine and femur datasets, MorphoFlow achieves high-fidelity reconstructions and accurately recovers population-consistent, structured patterns of anatomical variation.
This study addresses the challenge of effectively integrating structural and functional neuroimaging data by proposing a multimodal graph variational autoencoder (gMMVAE). The method introduces a modality-aware graph encoding mechanism that maps gray matter volume and static functional connectivity into a unified low-dimensional latent space. It further provides a systematic comparison of diverse generative architectures—including VAEs, Transformers, GANs, and diffusion models—in modeling graph-structured brain data. Experimental results demonstrate that gMMVAE consistently outperforms existing approaches in terms of generation fidelity, reconstruction quality, computational efficiency, and discriminative power of the latent representation. This work thus establishes an efficient and interpretable generative modeling paradigm for multimodal brain network analysis.
This study addresses the limitation of existing accelerated multi-contrast MRI reconstruction methods, which typically process each contrast independently and fail to fully exploit shared anatomical information across contrasts. To overcome this, we propose MAX, a framework that learns subject-specific anatomical representations via manifold expansion for efficient reconstruction. By employing intensity augmentation to expand the manifold, MAX effectively decouples shared anatomical structures from contrast-dependent components, integrating decoupled implicit neural representations with unrolled optimization techniques. Experimental results demonstrate that MAX achieves state-of-the-art PSNR and SSIM performance in both brain and knee MRI reconstruction, yielding improvements exceeding 1 dB. Furthermore, the proposed method exhibits significant robustness against motion artifacts and noise, highlighting its potential for reliable clinical application.
Existing self-supervised methods for ultrasound imaging often neglect anatomical context, hindering the learning of clinically aligned representations. This work proposes ANAUS, a novel framework that, for the first time, leverages anatomical structures as anchors in self-supervised learning. ANAUS introduces a learnable latent prompt engine combined with one-shot domain adaptation to enable annotation-free anatomical segmentation. It further incorporates a dual-strategy self-supervised mechanism—cross-view anatomical region semantic alignment and masked reconstruction of contextual core regions—to enhance representation invariance and fine-grained detail perception. Evaluated across six public datasets, ANAUS significantly outperforms state-of-the-art methods while maintaining computational efficiency suitable for clinical deployment.