🤖 AI Summary
Current AI-based clinical coding research is severely misaligned with real-world healthcare practice: mainstream evaluation focuses exclusively on the top-50 most frequent ICD codes, neglecting the long-tail distribution and complexity of thousands of infrequent but clinically critical codes. Method: Leveraging authentic U.S. electronic health record (EHR) data, this study conducts a methodological critique and human factors engineering analysis to systematically identify the root causes of evaluation bias in existing paradigms. Contribution/Results: We propose eight actionable recommendations to reform clinical coding evaluation—emphasizing code coverage, workflow integration, and practical utility—and reposition AI from end-to-end autonomous coding toward human-AI collaborative interaction that augments core coder tasks. Our work establishes a new, workflow-centered benchmark for clinical coding AI research, grounded in actual clinical practice and human-centered design principles.
📝 Abstract
Clinical coding is crucial for healthcare billing and data analysis. Manual clinical coding is labour-intensive and error-prone, which has motivated research towards full automation of the process. However, our analysis, based on US English electronic health records and automated coding research using these records, shows that widely used evaluation methods are not aligned with real clinical contexts. For example, evaluations that focus on the top 50 most common codes are an oversimplification, as there are thousands of codes used in practice. This position paper aims to align AI coding research more closely with practical challenges of clinical coding. Based on our analysis, we offer eight specific recommendations, suggesting ways to improve current evaluation methods. Additionally, we propose new AI-based methods beyond automated coding, suggesting alternative approaches to assist clinical coders in their workflows.