Evaluating the Generalization of Neuroimaging Foundation Models on African Brain MRI
研究评估了四个神经影像基础模型在非洲脑MRI数据上的泛化能力,发现这些模型在非西方小规模临床队列中的细粒度诊断分类效果不佳,而端到端训练的ViT3D表现更优。
研究评估了四个神经影像基础模型在非洲脑MRI数据上的泛化能力,发现这些模型在非西方小规模临床队列中的细粒度诊断分类效果不佳,而端到端训练的ViT3D表现更优。
We develop a network-structured Bayesian hierarchical model for sparse association mapping between genomic alterations and quantitative treatment-response phenotypes. The framework combines a Gaussian Markov random field prior that borrows strength across pathway-connected genes, a global-local horseshoe prior inducing sparsity, and a conjugate Gibbs sampler requiring no Metropolis-Hastings steps. Though broadly applicable to high-dimensional settings with known predictor networks, we validate it using cancer cell-line drug-sensitivity data. Applied to GDSC2 ($N=951$ cell lines, $G=219$ driver genes, $D=295$ drugs), the model identifies 126 gene-drug associations (0.195\% of 64{,}605 pairs), concentrated in EZH2 (45 drugs, all sensitivity-direction, mean effect $-0.911$ $\ln$IC50) and KMT2D (36 drugs, all sensitivity-direction, mean effect $-0.496$ $\ln$IC50). These markers show external support in an independent PRISM screen (1{,}518 compounds), with KMT2D achieving complete directional replication (36/36) and EZH2 partial replication (8/12). Five-fold cross-validated predictive log-likelihood confirms each prior layer's value: the full model outperforms the no-network ablation by $+3{,}109$ log-units per fold and the no-horseshoe ablation by $+14{,}039$ log-units, consistently across folds. Simulations under three scenarios show the full model achieves the highest precision and lowest false-discovery rate throughout, while the network prior improves sensitivity recovery under network-structured signal. A tissue-stratified extension identifies coherent subgroup refinements, including lung-specific EGFR-inhibitor sensitivity and skin-specific BRAF-Dabrafenib sensitivity. These results show the framework identifies sparse, interpretable, externally supported drug-sensitivity markers while enabling principled investigation of tissue-specific departures from shared effects.
This study addresses critical limitations in current automated chest X-ray report generation—namely, deficiencies in structural coherence, anatomical completeness, and semantic fidelity. To overcome these challenges, the authors propose a novel approach that integrates supervised fine-tuning of MedGemma-4B with a clinically grounded, rule-based procedural reward mechanism. Notably, this work introduces Group Relative Policy Optimization (GRPO) to medical image report generation for the first time, enabling interpretable alignment without reliance on neural reward models. The method synergistically combines a vision-language model with multidimensional rule-based rewards encompassing structural validation, anatomical checklist adherence, semantic similarity, and length constraints. In a blinded evaluation involving 69 cases, the system achieved an impression accuracy of 27.2% and a medical terminology usage rate of 86.5%, significantly outperforming both Gemini 2.5 Flash and the MedGemma-4B baseline.
Current medical benchmarks inadequately assess large language models’ ability to distinguish between clinically similar cases that require different interventions, often leading to an overestimation of their robustness. To address this gap, this work introduces MamaBench, the first counterfactual clinical diagnostic benchmark focused on maternal and child health, alongside an Evidence-Anchored Retrieval-Augmented Generation (EA-RAG) framework. EA-RAG employs a three-stage, evidence-coverage-driven retrieval mechanism—comprising clinical parameter extraction, coverage auditing, and contrastive sub-query generation—that overcomes the limitations of conventional similarity-based aggregation. Evaluated on Claude Sonnet 4.6, the approach achieves a robust accuracy of 65.0% and reduces the Bias Trap Rate (BTR) to 20.3%, a 5.5-percentage-point improvement over the baseline without compromising base performance, revealing that current models still fail in approximately 20% of counterfactual clinical reasoning scenarios.
This work addresses the critical gap in principled quantification of prediction uncertainty in clinical AI systems, which undermines their trustworthiness and fairness in high-stakes medical settings. The authors propose the first end-to-end Bayesian deep learning framework that integrates multimodal patient data through modality-specific variational encoders, precision-weighted late fusion, a composite Bayesian loss, and an uncertainty calibration penalty to disentangle aleatoric and epistemic uncertainties. Notably, they introduce calibrated uncertainty as a formal metric for algorithmic fairness and conduct cross-subgroup audits across demographic and socioeconomic strata. Experimental results demonstrate a low expected calibration error (ECE = 0.096) and reveal significant fairness gaps in uncertainty estimates for patients from primary/rural care settings, those with low socioeconomic status, and older adults (p < 0.001), while no significant disparity was observed across gender groups.
研究评估了四个神经影像基础模型在非洲脑MRI数据上的泛化能力,发现这些模型在非西方小规模临床队列中的细粒度诊断分类效果不佳,而端到端训练的ViT3D表现更优。
We develop a network-structured Bayesian hierarchical model for sparse association mapping between genomic alterations and quantitative treatment-response phenotypes. The framework combines a Gaussian Markov random field prior that borrows strength across pathway-connected genes, a global-local horseshoe prior inducing sparsity, and a conjugate Gibbs sampler requiring no Metropolis-Hastings steps. Though broadly applicable to high-dimensional settings with known predictor networks, we validate it using cancer cell-line drug-sensitivity data. Applied to GDSC2 ($N=951$ cell lines, $G=219$ driver genes, $D=295$ drugs), the model identifies 126 gene-drug associations (0.195\% of 64{,}605 pairs), concentrated in EZH2 (45 drugs, all sensitivity-direction, mean effect $-0.911$ $\ln$IC50) and KMT2D (36 drugs, all sensitivity-direction, mean effect $-0.496$ $\ln$IC50). These markers show external support in an independent PRISM screen (1{,}518 compounds), with KMT2D achieving complete directional replication (36/36) and EZH2 partial replication (8/12). Five-fold cross-validated predictive log-likelihood confirms each prior layer's value: the full model outperforms the no-network ablation by $+3{,}109$ log-units per fold and the no-horseshoe ablation by $+14{,}039$ log-units, consistently across folds. Simulations under three scenarios show the full model achieves the highest precision and lowest false-discovery rate throughout, while the network prior improves sensitivity recovery under network-structured signal. A tissue-stratified extension identifies coherent subgroup refinements, including lung-specific EGFR-inhibitor sensitivity and skin-specific BRAF-Dabrafenib sensitivity. These results show the framework identifies sparse, interpretable, externally supported drug-sensitivity markers while enabling principled investigation of tissue-specific departures from shared effects.
This study addresses critical limitations in current automated chest X-ray report generation—namely, deficiencies in structural coherence, anatomical completeness, and semantic fidelity. To overcome these challenges, the authors propose a novel approach that integrates supervised fine-tuning of MedGemma-4B with a clinically grounded, rule-based procedural reward mechanism. Notably, this work introduces Group Relative Policy Optimization (GRPO) to medical image report generation for the first time, enabling interpretable alignment without reliance on neural reward models. The method synergistically combines a vision-language model with multidimensional rule-based rewards encompassing structural validation, anatomical checklist adherence, semantic similarity, and length constraints. In a blinded evaluation involving 69 cases, the system achieved an impression accuracy of 27.2% and a medical terminology usage rate of 86.5%, significantly outperforming both Gemini 2.5 Flash and the MedGemma-4B baseline.
Current medical benchmarks inadequately assess large language models’ ability to distinguish between clinically similar cases that require different interventions, often leading to an overestimation of their robustness. To address this gap, this work introduces MamaBench, the first counterfactual clinical diagnostic benchmark focused on maternal and child health, alongside an Evidence-Anchored Retrieval-Augmented Generation (EA-RAG) framework. EA-RAG employs a three-stage, evidence-coverage-driven retrieval mechanism—comprising clinical parameter extraction, coverage auditing, and contrastive sub-query generation—that overcomes the limitations of conventional similarity-based aggregation. Evaluated on Claude Sonnet 4.6, the approach achieves a robust accuracy of 65.0% and reduces the Bias Trap Rate (BTR) to 20.3%, a 5.5-percentage-point improvement over the baseline without compromising base performance, revealing that current models still fail in approximately 20% of counterfactual clinical reasoning scenarios.
This work addresses the critical gap in principled quantification of prediction uncertainty in clinical AI systems, which undermines their trustworthiness and fairness in high-stakes medical settings. The authors propose the first end-to-end Bayesian deep learning framework that integrates multimodal patient data through modality-specific variational encoders, precision-weighted late fusion, a composite Bayesian loss, and an uncertainty calibration penalty to disentangle aleatoric and epistemic uncertainties. Notably, they introduce calibrated uncertainty as a formal metric for algorithmic fairness and conduct cross-subgroup audits across demographic and socioeconomic strata. Experimental results demonstrate a low expected calibration error (ECE = 0.096) and reveal significant fairness gaps in uncertainty estimates for patients from primary/rural care settings, those with low socioeconomic status, and older adults (p < 0.001), while no significant disparity was observed across gender groups.