đ€ AI Summary
This study addresses the challenge of structuring massive volumes of free-text radiology reports, where manual curation is prohibitively costly and archive-wide automation remains elusive. To this end, we propose the first unsupervised structuring framework leveraging a locally deployed, single-GPU open-source large language model (gpt-oss-120B). By integrating constrained decoding for three-level template selection with a hierarchical template repository and high-throughput inference, the approach automatically converts multimodal reports into structured data, thereby overcoming traditional dependencies on cloud computing and manual review. Evaluated on 2.18 million reports, the method achieves a 96.5% structuring rate with semantic similarity exceeding 0.95 and a processing throughput of 1,258 reports per hour, substantially reducing the need for human intervention.
đ Abstract
Purpose: To develop and evaluate an open-weight large language model (LLM) pipeline that converts an entire archive of free-text radiology reports into structured reports without human oversight. Materials and Methods: In this retrospective study, a pipeline with 150 hierarchically organized templates was developed at one center and tested at a second center on reports from 2010 to 2025. The open-weight model gpt-oss-120B selects the template in three constrained-decoding steps and fills it on one local graphics processing unit. Template selection was scored against expert labels on 914 randomly sampled reports of five modalities, structuring quality on 920 radiography and CT reports corrected field by field by five residents. The pipeline then processed the complete archive of the second center. Proportions are reported with Wilson 95% confidence intervals (CIs). Results: An optimal template set was selected for 74.4% of reports (680 of 914; 95% CI: 71.5%, 77.1%) and an appropriate set for 82.3% (752 of 914; 95% CI: 79.7%, 84.6%), 87.7% for single-region and 54.1% for multi-region reports. Macro semantic textual similarity between output and corrected reference was 0.95 for radiography and 0.97 for CT, residents left 88.7% of 24,638 fields unchanged, and unsupported content was flagged in 1.0% and 1.5% of reports. Of 2,186,982 archive reports, 96.5% received structured output, 2,401,544 structured reports, at 1,258 reports per hour on one graphics processing unit. Conclusion: An open-weight LLM pipeline structured a complete multimodality report archive without human oversight with high content fidelity. Multi-region reports remained the main source of template errors.