Hierarchical Divide-and-Conquer for Fine-Grained Alignment in LLM-Based Medical Evaluation

📅 2025-01-12
📈 Citations: 0
Influential: 0
📄 PDF

career value

170K/year
🤖 AI Summary
Existing clinical evaluation methods—such as multiple-choice questions—fail to capture the complexity of diagnostic and therapeutic decision-making, while current LLM evaluators exhibit insufficient fine-grained alignment with physician judgments. To address these limitations, we propose HDCEval, a novel clinical evaluation framework. First, we introduce a multidimensional, fine-grained assessment guideline co-developed by medical experts, covering three core dimensions: patient-problem relevance, medical knowledge correctness, and communication quality. Second, we propose Attribute-Driven Token Optimization (ADTO), a method that enables precise modeling of assessment signals via attribute-aware token-level optimization. Third, we design a hierarchical task decomposition scheme coupled with an expert-model collaborative discrimination mechanism. Evaluated on real-world clinical cases, HDCEval achieves significantly improved agreement with human physicians (Krippendorff’s α increased by +0.32) and consistently outperforms existing LLM evaluators, offering a more accurate reflection of clinical decision complexity and domain expertise.

Technology Category

Application Category

📝 Abstract
In the rapidly evolving landscape of large language models (LLMs) for medical applications, ensuring the reliability and accuracy of these models in clinical settings is paramount. Existing benchmarks often focus on fixed-format tasks like multiple-choice QA, which fail to capture the complexity of real-world clinical diagnostics. Moreover, traditional evaluation metrics and LLM-based evaluators struggle with misalignment, often providing oversimplified assessments that do not adequately reflect human judgment. To address these challenges, we introduce HDCEval, a Hierarchical Divide-and-Conquer Evaluation framework tailored for fine-grained alignment in medical evaluation. HDCEval is built on a set of fine-grained medical evaluation guidelines developed in collaboration with professional doctors, encompassing Patient Question Relevance, Medical Knowledge Correctness, and Expression. The framework decomposes complex evaluation tasks into specialized subtasks, each evaluated by expert models trained through Attribute-Driven Token Optimization (ADTO) on a meticulously curated preference dataset. This hierarchical approach ensures that each aspect of the evaluation is handled with expert precision, leading to a significant improvement in alignment with human evaluators.
Problem

Research questions and friction points this paper is trying to address.

Medical Assessment
Complexity of Clinical Practice
Evaluation Model Accuracy
Innovation

Methods, ideas, or system contributions that make the work stand out.

HDCEval
Medical Assessment
Large Language Model Evaluation
🔎 Similar Papers
No similar papers found.