PathReportEval: A Systematic Benchmark for Pathology Report Generation

📅 2026-07-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of clinically oriented evaluation criteria in pathological report generation, where conventional natural language generation metrics fail to detect critical errors such as missed diagnoses, hallucinations, or inconsistencies in tumor attributes. To bridge this gap, the authors establish a standardized benchmark and evaluation framework encompassing three datasets, four generation methods, and three pathology-specific visual encoders—CONCHv1.5, UNI2-h, and H-Optimus-1—with unified preprocessing, training, and evaluation protocols. The core contribution is the Clinical Report Quality Score (CRQS), an interpretable metric assessing reports along four clinical dimensions: clinical fact coverage, key information recall, hallucination rate, and clinical inconsistency. Experiments demonstrate that traditional metrics exhibit weak correlation with clinical correctness and often overestimate performance, whereas CRQS effectively reveals meaningful differences among models and encoders in terms of clinical fidelity.
📝 Abstract
Pathology report generation from whole-slide images (WSIs) is a rapidly growing multimodal learning problem, yet progress is difficult to measure because existing studies use heterogeneous datasets, model settings, visual encoders, and evaluation protocols. Moreover, commonly used natural language generation metrics, including BLEU, ROUGE, and METEOR, primarily reward lexical similarity and often fail to detect clinically consequential errors such as omitted diagnoses, hallucinated findings, or discordant tumor attributes. We present a standardized benchmark and evaluation framework for pathology report generation. The benchmark evaluates four representative methods across three datasets (TCGA, HistAI, and REG 2025) using three pathology foundation encoders (CONCHv1.5, UNI2-h, and H-Optimus-1). Our framework standardizes preprocessing, feature extraction, training, decoding, and evaluation, enabling fair comparison across models while providing a modular platform for integrating new methods, datasets, and encoders. A central contribution is the Clinical Report Quality Score (CRQS), a clinically grounded metric for evaluating factual correctness. CRQS maps reference and generated reports into structured clinical attributes and measures four complementary dimensions: clinical fact coverage, key information recall, hallucination rate, and clinical discordance, producing both an overall score and interpretable sub-scores. Experiments demonstrate that conventional language-generation metrics are weakly aligned with clinical correctness and frequently overestimate report quality. In contrast, CRQS reveals clinically meaningful differences between models and encoders that lexical metrics fail to capture. Together, the benchmark, public plug-and-play framework, and CRQS establish a reproducible foundation for rigorous evaluation of pathology report generation.
Problem

Research questions and friction points this paper is trying to address.

pathology report generation
evaluation benchmark
clinical correctness
natural language generation metrics
whole-slide images
Innovation

Methods, ideas, or system contributions that make the work stand out.

pathology report generation
standardized benchmark
Clinical Report Quality Score (CRQS)
multimodal learning
clinical evaluation metrics
🔎 Similar Papers
No similar papers found.