🤖 AI Summary
This study addresses the security risks in large language models (LLMs) performing rule-based decision-making tasks, where incomplete reasoning rationales remain difficult to detect. To overcome the inability of existing evaluators to effectively identify missing justifications, this work constructs a dedicated decision-making and reasoning benchmark and proposes a verifiable rationale completeness checking mechanism that integrates label-blind extraction, deterministic verification, and reward model training. Experimental results demonstrate that the proposed method achieves an AUROC of 69.24%, outperforming baseline approaches by 10.37 percentage points and significantly surpassing existing evaluators. This advancement effectively fills a critical gap in the field of rationale consistency detection for LLMs.
📝 Abstract
We study reasoning in Rule-Governed Decision Tasks (RGDTs), where models apply external rules to case facts and justify decisions, as required in policy, contract, and compliance settings. Beyond the deductive capability emphasized by standard mathematical and logical reasoning tasks, RGDTs require interpreting rules and their applicability, assessing conditions from evidence, combining judgments under rules and exceptions, and providing checkable justifications. These demands motivate a benchmark assessing both decisions and their stated grounds. We introduce RGDT-Bench, providing 202.1K condition-level supervision slots across four task tracks and eight supported task-probe combinations that vary access to supporting information. Label-blind extraction and deterministic checks produce labels for warrant completeness: source-referenced coverage and consistency of stated decision grounds. The benchmark attributes failures to four process layers: rule use, condition, evidence, and aggregation, and checks the final outcome. Among evaluable correct responses, warrant incompleteness averages 40.2% across six evaluated LLMs and supported task-probe combinations. Such warrant incompleteness poses potential safety risks and remains difficult to detect: the best of seventeen existing evaluators reaches only 57.69% (random: 50%) task-averaged area under the receiver operating characteristic curve (AUROC). To address this difficulty, we train a simple reward model with warrant supervision. It achieves 69.24% task-averaged AUROC among correct answers, exceeding the matched outcome-supervised baseline by 10.37 pp (percentage points) and the best existing evaluator by 11.55 pp. Beyond completeness assessment, the model outperforms both outcome-supervised baselines across nearly all response-selection comparisons, supporting RGDT-Bench's warrant supervision for RGDT reasoning.