🤖 AI Summary
This study addresses the challenge of misidentifying critical evidence in rule-based reasoning, which leads to erroneous decisions, by proposing the InterPact framework. This method introduces an evidence intervention generator and a propagation verifier grounded in counterfactual intervention learning. By leveraging complete state-to-decision mapping supervision, it directly assesses evidence criticality without relying on external powerful models. Furthermore, it integrates fact editing, compositional operations, and conditional probability-weighted prediction with frozen language models. Experimental results demonstrate that InterPact achieves 68.28% accuracy on single-case evidence criticality verification tasks, significantly outperforming all baselines. This effectively supports the prioritized review of sensitive evidence, thereby enhancing the safety of model decision-making.
📝 Abstract
Rule-based reasoning, as in eligibility checks and contract reviews, requires language models to assess evidence against individual conditions and combine their judgments under explicit rules. Errors in evidence assessment can leave a decision unchanged, but misinterpreting or overlooking decision-critical evidence can reverse it. Identifying such evidence allows more capable models to focus on checking the corresponding condition judgments, supporting accurate and safe decisions. Recognizing the evidence's criticality requires understanding how evidence affects a condition judgment and how that judgment affects the decision. To achieve the goal, we propose a INTERvention-based imPACT learning framework (InterPact), which enables counterfactual verification of evidence criticality in rule-governed decisions. Specifically, its evidence intervention constructor generates training pairs for a propagation verifier by editing case facts with a frozen language model while holding rules and non-target conditions fixed. Human-reviewed labels record the resulting condition and decision changes, while complete state-to-decision mappings supervise consequences beyond the observed edit. During training, the verifier weights learned conditional decision predictions by evidence-based condition probabilities through a fixed composition operation, propagating decision-change supervision into the base model. At inference, the trained base model directly judges criticality from the original case and target evidence, without human or stronger-model supervision. On single-case evidence criticality verification over adapted rule-governed decision cases, InterPact achieves 68.28% accuracy, outperforming all six baselines. These results support learned decision sensitivity as a basis for prioritizing evidence checks.