🤖 AI Summary
This study addresses the privacy, cost, and reproducibility bottlenecks of proprietary models in radiology report entity extraction by exploring open-source alternatives. Methodologically, we employ the Gemma-3-12B model, integrating a discriminative classification head with generative instruction tuning, and construct training data through knowledge distillation from authentic clinical reports. Our findings demonstrate that data provenance is critical: the model distilled from real reports achieves a macro-F1 score of 0.845, outperforming the baseline by 0.178 and attaining performance comparable to GPT-4o without statistically significant differences. This work enables high-performance local deployment for intracranial hemorrhage label extraction on consumer-grade GPUs, establishing a low-cost, privacy-preserving open-source paradigm for medical artificial intelligence.
📝 Abstract
Converting free-text radiology reports into structured labels supports cohort building, quality assurance, and monitoring of clinical imaging models, but the strongest label extractors are hosted proprietary models whose use raises privacy, cost, and reproducibility concerns. We asked whether a fine-tuned open-weight model (Gemma-3-12B) can match GPT-4o at multi-label intracranial hemorrhage (ICH) acuity extraction from non-contrast head-CT reports, and which ingredients matter. Using a 2x2 design, we crossed two adaptation strategies (a discriminative classification head, CH; generative instruction fine-tuning, IFT) with two training-data sources (distillation of real GPT-4o-labeled reports; synthetic reports generated by GPT-4o from real exemplars), across five training sizes, benchmarked on 100 expert-adjudicated reports against GPT-4o and the un-tuned open-weight base. The distilled instruction-tuned model (DIFT) matched GPT-4o (macro-F1 0.845 vs 0.850; p = 1.000) and exceeded the base model by 0.178. The decisive factor was the training-data source, not the fine-tuning method: both synthetic-data models failed to exceed the un-tuned open-weight base at any training size and underperformed the distilled models across all acuity classes. Fine-tuning and inference fit within the memory envelope of a single 24 GB consumer GPU. For narrow, high-value clinical label-extraction tasks, distilling real reports, rather than generating synthetic ones, is what closes the gap to a hosted model, enabling a private, low-cost, version-stable on-premises alternative.