Beyond Literal Summarization: Redefining Hallucination for Medical SOAP Note Evaluation
Current evaluation methods for clinical document generation often misclassify valid medical reasoning as hallucination, leading to an underestimation of large language model performance. This work proposes a clinically grounded evaluation framework that redefines “hallucination” in SOAP notes—not by literal fidelity but by clinical plausibility—through calibrated prompt engineering, a retrieval-augmented mechanism supported by medical ontologies such as SNOMED CT, and a reasoning-aware assessment pipeline. The framework effectively distinguishes genuine hallucinations from legitimate clinical abstractions, including terminology mapping, diagnostic inference, and guideline-concordant care planning. Experimental results demonstrate a significant reduction in average hallucination rates from 35% to 9%, with the remaining cases predominantly involving actual safety concerns, thereby validating the framework’s efficacy and clinical relevance.