LearnActCoder: Role-Aware Error Memory for Adaptive Clinical Coding Agents
本文提出Learn-Then-Act框架,通过构建结构化错误知识库来解决临床编码中的重复性错误问题,提高了CPT编码的F1分数。
本文提出Learn-Then-Act框架,通过构建结构化错误知识库来解决临床编码中的重复性错误问题,提高了CPT编码的F1分数。
本文提出了一种结构化证据路由方法,通过路由器-预测器-审查者工作流程解决从多模态纵向电子健康记录中预测事件风险的问题。
研究通过结合结构化电子健康记录基础模型与语言基础模型,并使用ICD-Deepresearch流程,提高了未来就诊时诊断代码预测的准确性和医生对检索文档的认可度。
Existing approaches to filtering high-quality reasoning data rely heavily on strong reasoning models, resulting in high costs and limited effectiveness. This work proposes an efficient alternative that reliably identifies challenging samples by analyzing the loss of a pretrained model over the first 100 reasoning tokens. By further incorporating loss patterns and gradient similarity from a small number of perturbed checkpoints, the method enables precise selection of diverse, high-difficulty data without requiring complex inference procedures. Evaluated on Qwen2.5-7B and Llama3.1-8B, this approach achieves up to a 1.7% improvement in fine-tuning performance while reducing token consumption by 91%, substantially lowering the cost of data curation.
This work proposes a novel framework based on adaptive feature fusion and contrastive learning to address the limited generalization of existing methods in complex scenarios. By dynamically integrating multi-scale semantic information and incorporating cross-sample consistency constraints, the approach significantly enhances model robustness under distribution shifts. Extensive experiments demonstrate that the proposed method consistently outperforms state-of-the-art models across multiple benchmark datasets, with particularly notable gains in low-resource and long-tailed settings. Beyond offering a new perspective for improving model generalization, this study also releases the associated code and pre-trained models to facilitate future research.
本文提出Learn-Then-Act框架,通过构建结构化错误知识库来解决临床编码中的重复性错误问题,提高了CPT编码的F1分数。
本文提出了一种结构化证据路由方法,通过路由器-预测器-审查者工作流程解决从多模态纵向电子健康记录中预测事件风险的问题。
研究通过结合结构化电子健康记录基础模型与语言基础模型,并使用ICD-Deepresearch流程,提高了未来就诊时诊断代码预测的准确性和医生对检索文档的认可度。
Existing approaches to filtering high-quality reasoning data rely heavily on strong reasoning models, resulting in high costs and limited effectiveness. This work proposes an efficient alternative that reliably identifies challenging samples by analyzing the loss of a pretrained model over the first 100 reasoning tokens. By further incorporating loss patterns and gradient similarity from a small number of perturbed checkpoints, the method enables precise selection of diverse, high-difficulty data without requiring complex inference procedures. Evaluated on Qwen2.5-7B and Llama3.1-8B, this approach achieves up to a 1.7% improvement in fine-tuning performance while reducing token consumption by 91%, substantially lowering the cost of data curation.
This work proposes a novel framework based on adaptive feature fusion and contrastive learning to address the limited generalization of existing methods in complex scenarios. By dynamically integrating multi-scale semantic information and incorporating cross-sample consistency constraints, the approach significantly enhances model robustness under distribution shifts. Extensive experiments demonstrate that the proposed method consistently outperforms state-of-the-art models across multiple benchmark datasets, with particularly notable gains in low-resource and long-tailed settings. Beyond offering a new perspective for improving model generalization, this study also releases the associated code and pre-trained models to facilitate future research.