Score
Designs and implements rules, mappings, and encoded representations that translate clinical information—such as diagnoses, procedures, and laboratory tests—into standard clinical codes. Ensures preservation of hierarchical code relationships and aligns code selection and frequency distributions with cohort definitions for accurate reporting, billing, analysis, and interoperability.
Current AI-based clinical coding research is severely misaligned with real-world healthcare practice: mainstream evaluation focuses exclusively on the top-50 most frequent ICD codes, neglecting the long-tail distribution and complexity of thousands of infrequent but clinically critical codes. Method: Leveraging authentic U.S. electronic health record (EHR) data, this study conducts a methodological critique and human factors engineering analysis to systematically identify the root causes of evaluation bias in existing paradigms. Contribution/Results: We propose eight actionable recommendations to reform clinical coding evaluation—emphasizing code coverage, workflow integration, and practical utility—and reposition AI from end-to-end autonomous coding toward human-AI collaborative interaction that augments core coder tasks. Our work establishes a new, workflow-centered benchmark for clinical coding AI research, grounded in actual clinical practice and human-centered design principles.
Current general-purpose large language models (LLMs) suffer from hierarchical misalignment—generating semantically proximal yet incorrect ICD codes—in clinical coding tasks. Moreover, mainstream benchmarks (e.g., MIMIC) exhibit critical limitations: insufficient supporting evidence in notes and strong inpatient bias, undermining model reliability and generalizability. To address these issues, we propose a trustworthy clinical coding framework comprising three core components: (1) formalizing code-level hierarchical validation as a novel auxiliary task; (2) constructing the first expert-annotated, multi-department outpatient benchmark dataset with dual annotations per case; and (3) integrating prompt engineering, lightweight fine-tuning, and an ICD hierarchy-aware verification module for end-to-end error correction. Experiments demonstrate substantial reductions in near-miss errors, significant improvements in coding accuracy (+4.2% macro-F1) and robustness across diverse clinical scenarios. Our framework establishes a high-fidelity, clinically grounded solution for automated ICD coding.
This study addresses the inefficiency and systematic undercoding of secondary diagnoses in manual medical coding by developing a multimodal language model trained on 5.8 million electronic health records from 1.8 million patients in eastern Denmark—a population-scale cohort encompassing nearly all medical specialties. The model integrates clinical notes, medication records, and laboratory data to predict ICD-10 codes. Evaluated on a hold-out set of 270,000 patients, it achieves a micro-averaged F1 score of 71.8% and a top-10 recall of 95.5%. It also identified thousands of cases with missed secondary diagnoses, 76–86% of which were confirmed as valid upon manual review. The approach can automate approximately 50% of coding tasks, offering a scalable tool for epidemiological and multimorbidity research.
This work addresses the limitation of existing electronic health record (EHR) foundation models that treat ICD diagnosis codes as flat tokens, thereby neglecting their intrinsic clinical hierarchical structure and underutilizing semantic information. To overcome this, the study introduces the ICD-10-CM hierarchy as an inductive bias into EHR modeling for the first time, proposing a hierarchy-enhanced Transformer and a hierarchy-aware graph neural network that jointly leverage diagnosis co-occurrence patterns and multi-granular ICD codes. The model is pretrained on MIMIC-IV and evaluated via frozen probing on eICU for cross-dataset generalization. Results demonstrate consistent and significant improvements over flat-code baselines in both in-domain and cross-domain settings, enhancing downstream prediction performance and yielding more semantically coherent embedding spaces, thus validating the broad utility of hierarchical modeling across diverse tasks and architectures.
Inconsistent diagnostic coding and poor interoperability across institutions and species hinder effective utilization of veterinary electronic health records (EHRs). Method: We propose an automated SNOMED-CT diagnostic coding framework leveraging large language models (LLMs), fine-tuning ten open-source Transformer architectures on 246,000 manually annotated clinical notes from the Colorado State University Veterinary Teaching Hospital. Contribution/Results: To our knowledge, this is the first approach achieving full coverage mapping to all 7,739 SNOMED-CT diagnosis codes used clinically in that institution. The best-performing model achieves an F1-score of 0.82—significantly outperforming baselines such as DeepTag and VetTag. Notably, even non-clinically pre-trained LLMs attain F1 > 0.78 under limited annotation budgets, demonstrating robust generalizability and feasibility in resource-constrained settings. This work establishes a scalable, low-cost paradigm for interoperable, cross-institutional integration of veterinary health data.
This study addresses the frequent violations of clinical coding standards—such as ICD-10, CPT, and HL7 FHIR—by large language models when generating structured medical data, which impedes integration with electronic health record systems. To mitigate this, the authors propose and validate a closed-loop verification-and-repair framework that automatically detects and iteratively corrects formatting errors. The approach is evaluated using three open-source models—Qwen2.5-7B, Llama3.1-8B, and Gemma2-9B—deployed locally across 320 clinical scenarios. Results demonstrate a substantial improvement in schema compliance across all models, achieving an overall adherence rate of 99.0% and increasing individual model performance by 7.8 to 12.5 percentage points. Notably, 96% of detected errors were attributable to repairable representation-layer issues, with most resolved within one or two correction rounds, effectively compensating for the models’ limited understanding of healthcare IT standards.
This work addresses the longstanding reliance on manual medical coding, which is inefficient and error-prone, and overcomes key limitations of existing automated approaches—namely poor generalizability to new coding systems and lack of interpretability. The authors propose an agent-based reasoning framework that emulates human expert decision-making by dynamically retrieving official coding guidelines and integrating them with clinical text understanding, enabling adaptation to any coding system without retraining. This approach achieves zero-shot cross-system transfer—the first of its kind—and provides traceable justifications linking predicted codes to supporting evidence in the source documents. Evaluated across five real-world and publicly available datasets spanning multiple countries and clinical specialties, the method demonstrates state-of-the-art performance and strong practical deployability.
This work addresses the challenge of constructing interpretable patient representations from multiple ICD diagnosis codes while maintaining high predictive performance. The authors propose a similarity-based relative grouping mechanism that integrates clinical diagnostic groupings with semantic information from pretrained ICD embeddings, such as ICD2Vec. By applying similarity-weighted mapping, high-dimensional diagnosis codes are projected into a low-dimensional feature space where each dimension corresponds to a clinically meaningful category. This approach preserves strong representational capacity while enhancing model interpretability through explicit clinical semantics. Extensive experiments on multiple large-scale electronic health record datasets demonstrate that the proposed method achieves predictive performance comparable to purely embedding-based approaches, yet offers substantially improved clinical interpretability of the resulting features.
This study addresses the clinical challenges of complex Evaluation and Management (E/M) coding, high manual annotation burden, and low billing efficiency. We propose ProFees—a large language model (LLM)-based framework employing multi-step reasoning and structured prompting for CPT coding. Unlike single-step prompting or opaque commercial systems, ProFees explicitly models the multidimensional E/M rules (e.g., history, physical examination, medical decision-making), enabling interpretable and verifiable automated coding. Evaluated on an expert-annotated dataset of real-world clinical documentation, ProFees achieves 89.2% coding accuracy—outperforming leading commercial systems by 36.1% and the best single-prompt baseline by 4.8%. To our knowledge, this is the first work to systematically introduce a structured multi-step reasoning paradigm to E/M coding, significantly improving accuracy, robustness, and clinical trustworthiness.
This work addresses the challenge of ensuring correctness and reliability in code that translates structured data within the Internet of Medical Things. The authors propose a novel code generation approach that integrates large language models with evolutionary algorithms and, for the first time, embeds formal verification into the LLM-driven synthesis pipeline. This integration guarantees that the generated JSON-to-FHIR transformation code strictly adheres to predefined specifications. The method significantly enhances translation accuracy and safety in healthcare data interoperability, as demonstrated in a pulse oximeter integration scenario. It enables the cost-effective and stable production of FHIR-compliant, reliable code, thereby advancing trustworthy medical device connectivity and data exchange.