π€ AI Summary
This study investigates the reliability and error propagation of large language models (LLMs) across the full decision chain in insurance claims processing. We construct an end-to-end evaluation benchmark spanning multiple insurance types and propose a cross-level controlled factual perturbation method. By integrating structured rule matching with multi-level consistency analysis, we systematically evaluate the modelβs complete reasoning pipeline from atomic rule parsing to payout calculation. Our findings reveal that LLM reliability degrades significantly along the decision chain; local step accuracy does not guarantee global decision correctness, and joint accuracy remains substantially lower than single-step metrics. This work elucidates the error accumulation mechanisms of LLMs in complex business scenarios, providing critical insights for enhancing their decision-making trustworthiness.
π Abstract
Recent advances in reasoning-oriented large language models (LLMs) have motivated increasing evaluation of their ability to perform professional decision tasks. Insurance claim adjudication is one such task, requiring models to connect case evidence, insurance rules, intermediate judgments, and payout calculations across a structured decision process. We introduce InsClaimBench, an end-to-end benchmark for evaluating insurance claim adjudication across the decision chain. Grounded in real claim materials and structured insurance rules, InsClaimBench contains 3,780 cases in 375 case families across auto, property, and health insurance, comprising 86,656 atomic rule judgments. It evaluates each claim from atomic rules through adjudication modules to payout decisions and amounts, with controlled factual variants testing whether required changes are correctly propagated across levels. Evaluation of six LLMs reveals a progressive loss of reliability along the decision chain. Payout-decision accuracy ranges from 74.23--80.19%, while joint decision--amount accuracy drops to 47.54--73.15%. Strong local performance also fails to ensure case-level correctness: atomic-rule accuracy reaches 95.48%, whereas rule-vector exact match peaks at only 36.90%, and the most frequent module errors are not necessarily those most associated with final-decision failure. Under factual changes, these inconsistencies further become propagation failures: module updates are less reliable than rule updates, correct local judgments can still yield incorrect payouts, and correct payouts can conceal intermediate errors. These results show that reliable claim adjudication requires consistent composition and propagation across the decision chain.