InsClaimBench: Benchmarking Insurance Claim Adjudication Across the Decision Chain

πŸ“… 2026-10-07
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study investigates the reliability and error propagation of large language models (LLMs) across the full decision chain in insurance claims processing. We construct an end-to-end evaluation benchmark spanning multiple insurance types and propose a cross-level controlled factual perturbation method. By integrating structured rule matching with multi-level consistency analysis, we systematically evaluate the model’s complete reasoning pipeline from atomic rule parsing to payout calculation. Our findings reveal that LLM reliability degrades significantly along the decision chain; local step accuracy does not guarantee global decision correctness, and joint accuracy remains substantially lower than single-step metrics. This work elucidates the error accumulation mechanisms of LLMs in complex business scenarios, providing critical insights for enhancing their decision-making trustworthiness.
πŸ“ Abstract
Recent advances in reasoning-oriented large language models (LLMs) have motivated increasing evaluation of their ability to perform professional decision tasks. Insurance claim adjudication is one such task, requiring models to connect case evidence, insurance rules, intermediate judgments, and payout calculations across a structured decision process. We introduce InsClaimBench, an end-to-end benchmark for evaluating insurance claim adjudication across the decision chain. Grounded in real claim materials and structured insurance rules, InsClaimBench contains 3,780 cases in 375 case families across auto, property, and health insurance, comprising 86,656 atomic rule judgments. It evaluates each claim from atomic rules through adjudication modules to payout decisions and amounts, with controlled factual variants testing whether required changes are correctly propagated across levels. Evaluation of six LLMs reveals a progressive loss of reliability along the decision chain. Payout-decision accuracy ranges from 74.23--80.19%, while joint decision--amount accuracy drops to 47.54--73.15%. Strong local performance also fails to ensure case-level correctness: atomic-rule accuracy reaches 95.48%, whereas rule-vector exact match peaks at only 36.90%, and the most frequent module errors are not necessarily those most associated with final-decision failure. Under factual changes, these inconsistencies further become propagation failures: module updates are less reliable than rule updates, correct local judgments can still yield incorrect payouts, and correct payouts can conceal intermediate errors. These results show that reliable claim adjudication requires consistent composition and propagation across the decision chain.
Problem

Research questions and friction points this paper is trying to address.

Insurance claim adjudication
Large language models
Decision chain
Benchmark evaluation
Error propagation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Insurance Claim Adjudication
Decision Chain Benchmark
Large Language Models
Error Propagation
Atomic Rule Evaluation
Linqi Zhang
Linqi Zhang
Professor of Microbiology and Immunology, Tsinghua University
virologyimmunologyantibodyvaccines
C
Chong Qi
School of Integrated Circuits, Nanjing University
Y
Yan Cheng
School of Economics, Fudan University
W
Wanqing Cao
School of Economics, Fudan University
Y
Yu Liu
School of Computer Science and Technology, Fudan University
C
Chenwei Lin
School of Computer Science and Technology, Fudan University
Xian Xu
Xian Xu
Fudan University
insurancedisaster economics